什么造就了好的智能体数据?——面向大语言模型智能体数据生成的ACE视角
What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
August 27, 2026
作者: Xingshan Zeng, Zishan Xu, Boju Zhang, Yuzhou Wu, Lingzhi Wang, Jianghao Lin, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Weinan Zhang, Yong Yu, Qun Liu, Weiwen Liu
cs.AI
摘要
LLM智能体越来越依赖生成的交互数据来学习如何与外部环境交互。智能体数据生成必须在环境、任务、交互和成功信号之间保持一致性,同时产生有用而非仅仅丰富多样的经验。现有研究覆盖了多个智能体领域,但以领域为中心的组织方式和异构评估往往掩盖了共同的生成机制,并将候选构建与验证和选择混为一谈。本研究为该领域构建了一个两层框架。首先,我们将智能体数据表示为一个通用的因子化对象(E,q,τ,v),包含环境规范、任务信号、交互实现和可选的验证器。我们按照主要锚点和依赖结构来组织生成范式。其次,我们通过准确性-复杂性-多样性(ACE)视角将生成形式化为约束分布设计。准确性建立了基于真实环境且内部一致的数据的可行支撑集。在该支撑集内,复杂性根据所声明学习器及其执行配置的能力来分配学习质量,而多样性则控制数据的覆盖范围和冗余度。利用该框架,我们探讨了先前工作如何验证生成的经验、构建和校准难度以及扩展行为覆盖范围。文献揭示了一种向执行锚定的准确性、学习器相对的复杂性和超越表面变化或数据集规模的多样性转变的趋势。我们进一步通过ACE视角讨论了智能体数据生成中更广泛的方向和新兴趋势,包括它们对扩展、数据来源、训练范式和自适应学习的影响。总体而言,核心挑战不仅仅是生成更多数据,而是在智能体和环境不断演变的过程中持续分配有效、信息丰富且不冗余的经验。
English
LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agentic data generation must maintain consistency among environments, tasks, interactions, and success signals while producing experience that is useful rather than merely abundant. Existing work spans many agent domains, but domain-centered organization and heterogeneous evaluation often obscure common generation mechanisms and conflate candidate construction with verification and selection. This work develops a two-level framework for the field. First, we represent agentic data as a common factorized object (E,q,τ,v), comprising an environment specification, task signal, interaction realization, and optional verifier. We organize generation paradigms by their primary anchor and dependency structure. Second, we formulate generation as constrained distribution design through the Accuracy-Complexity-divErsity (ACE) lens. Accuracy establishes the feasible support of grounded and internally consistent data. Within this support, Complexity places learning mass relative to the capability of a declared learner and execution configuration, while divErsity controls coverage and redundancy of data. Using this framework, we explore how prior work verifies generated experience, constructs and calibrates difficulty, and expands behavioral coverage. The literature reveals a shift toward execution-grounded accuracy, learner-relative complexity, and diversity beyond surface variation or dataset size. We further discuss broader directions and emerging trends in agentic data generation through the ACE lens, including their implications for scaling, data sources, training regimes and adaptive learning. Overall, the central challenge is not simply to generate more data, but to continually allocate valid, informative, and non-redundant experience as agents and environments evolve.