ChatPaper.aiChatPaper

優れたエージェント的データとは何か?—LLMエージェント向けデータ生成へのACEレンズ—

What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents

August 27, 2026
著者: Xingshan Zeng, Zishan Xu, Boju Zhang, Yuzhou Wu, Lingzhi Wang, Jianghao Lin, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Weinan Zhang, Yong Yu, Qun Liu, Weiwen Liu
cs.AI

要旨

LLMエージェントは、外部環境との相互作用方法を学習するために、生成された相互作用データに依存することが増えている。エージェント的データ生成(agentic data generation)は、環境、タスク、相互作用、成功シグナルの間で一貫性を維持しつつ、単に豊富であるだけでなく有用な経験を生み出さなければならない。既存の研究は多くのエージェント領域に及んでいるが、領域中心の整理と不均質な評価は、共通の生成メカニズムを曖昧にし、候補構築と検証・選択を混同することが多い。本稿では、この分野に対して2層の枠組みを構築する。第一に、エージェント的データを共通の分解されたオブジェクト(E,q,τ,v)として表現する。これは、環境仕様、タスク信号、相互作用の実現、任意の検証器から構成される。我々は、生成パラダイムをその主要なアンカーと依存構造によって整理する。第二に、生成を、正確性-複雑性-多様性(Accuracy-Complexity-divErsity: ACE)の観点を通した制約付き分布設計として定式化する。正確性は、接地され内部整合性のあるデータの実行可能な台集合を確立する。この台集合内において、複雑性は、宣言された学習器と実行構成の能力に相対する学習質量を配置し、多様性はデータのカバレッジと冗長性を制御する。この枠組みを用いて、先行研究が生成された経験をどのように検証し、難易度を構築・較正し、行動のカバレッジを拡張しているかを考察する。文献レビューは、実行に接地された正確性、学習器に相対的な複雑性、および表面的な変動やデータセットサイズを超えた多様性への移行を示している。さらに、ACEの観点を通してエージェント的データ生成におけるより広範な方向性と新たなトレンドを議論し、スケーリング、データソース、訓練レジーム、適応的学習への影響を含める。全体として、中心的な課題は単により多くのデータを生成することではなく、エージェントと環境が進化するにつれて、有効で情報量が多く非冗長な経験を継続的に割り当てることである。
English
LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agentic data generation must maintain consistency among environments, tasks, interactions, and success signals while producing experience that is useful rather than merely abundant. Existing work spans many agent domains, but domain-centered organization and heterogeneous evaluation often obscure common generation mechanisms and conflate candidate construction with verification and selection. This work develops a two-level framework for the field. First, we represent agentic data as a common factorized object (E,q,τ,v), comprising an environment specification, task signal, interaction realization, and optional verifier. We organize generation paradigms by their primary anchor and dependency structure. Second, we formulate generation as constrained distribution design through the Accuracy-Complexity-divErsity (ACE) lens. Accuracy establishes the feasible support of grounded and internally consistent data. Within this support, Complexity places learning mass relative to the capability of a declared learner and execution configuration, while divErsity controls coverage and redundancy of data. Using this framework, we explore how prior work verifies generated experience, constructs and calibrates difficulty, and expands behavioral coverage. The literature reveals a shift toward execution-grounded accuracy, learner-relative complexity, and diversity beyond surface variation or dataset size. We further discuss broader directions and emerging trends in agentic data generation through the ACE lens, including their implications for scaling, data sources, training regimes and adaptive learning. Overall, the central challenge is not simply to generate more data, but to continually allocate valid, informative, and non-redundant experience as agents and environments evolve.