什麼造就優質的智能體資料?以ACE視角審視LLM智能體的資料生成
What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
August 27, 2026
作者: Xingshan Zeng, Zishan Xu, Boju Zhang, Yuzhou Wu, Lingzhi Wang, Jianghao Lin, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Weinan Zhang, Yong Yu, Qun Liu, Weiwen Liu
cs.AI
摘要
LLM代理日益依賴於生成的互動資料來學習如何與外部環境互動。代理性資料生成必須在環境、任務、互動與成功訊號之間維持一致性,同時產出有用而非僅是大量的經驗。現有研究涵蓋眾多代理領域,但以領域為中心的組織方式與異質化的評估標準,往往掩蓋了共同的生成機制,並將候選建構與驗證和篩選混為一談。本研究為該領域建立了一個雙層框架。首先,我們將代理性資料表示為一個共同的分解物件(E,q,τ,v),包含環境規格、任務訊號、互動實現及可選的驗證器。我們依其主要錨點與依賴結構來組織生成範式。其次,我們透過「準確性-複雜性-多樣性」(ACE)視角,將生成形式化為約束分佈設計。準確性確立了具接地性且內部一致的資料之可行支撐範圍。在此支撐範圍內,複雜性依據所宣告的學習者能力與執行配置來配置學習質量,而多樣性則控制資料的覆蓋度與冗餘度。藉由該框架,我們探討既有研究如何驗證生成的經驗、如何建構並校準難度,以及如何擴展行為覆蓋範圍。文獻呈現出一種轉向:朝向以執行為基礎的準確性、以學習者為相對基準的複雜性,以及超越表面變化或資料集規模的多樣性。我們進一步透過ACE視角討論代理性資料生成的更廣泛方向與新興趨勢,包括其對規模化、資料來源、訓練機制與自適應學習的啟示。總體而言,核心挑戰不在於單純生成更多資料,而在於隨著代理與環境的演進,持續配置有效、具資訊性且不冗餘的經驗。
English
LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agentic data generation must maintain consistency among environments, tasks, interactions, and success signals while producing experience that is useful rather than merely abundant. Existing work spans many agent domains, but domain-centered organization and heterogeneous evaluation often obscure common generation mechanisms and conflate candidate construction with verification and selection. This work develops a two-level framework for the field. First, we represent agentic data as a common factorized object (E,q,τ,v), comprising an environment specification, task signal, interaction realization, and optional verifier. We organize generation paradigms by their primary anchor and dependency structure. Second, we formulate generation as constrained distribution design through the Accuracy-Complexity-divErsity (ACE) lens. Accuracy establishes the feasible support of grounded and internally consistent data. Within this support, Complexity places learning mass relative to the capability of a declared learner and execution configuration, while divErsity controls coverage and redundancy of data. Using this framework, we explore how prior work verifies generated experience, constructs and calibrates difficulty, and expands behavioral coverage. The literature reveals a shift toward execution-grounded accuracy, learner-relative complexity, and diversity beyond surface variation or dataset size. We further discuss broader directions and emerging trends in agentic data generation through the ACE lens, including their implications for scaling, data sources, training regimes and adaptive learning. Overall, the central challenge is not simply to generate more data, but to continually allocate valid, informative, and non-redundant experience as agents and environments evolve.