エージェント的アーティファクト生成:システム、評価、原理、および機会

Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities

August 28, 2026
著者: Tianfu Wang, Zhezheng Hao, Xilin Xia, Lixin Liu, Mengkang Hu, Hongzhang Liu, Xi Chen, Ziyan Liu, Xiankun Lin, Weijia Zhang, Nicholas Jing Yuan, Hui Xiong
cs.AI

要旨

生成モデルは、自然言語プロンプトを画像、テキスト、コード、その他のコンテンツに変換し、ドラフトや構成要素の作成コストを低減できる。その実用的な影響は、それらの断片が完全で信頼性のある成果物になり得るかどうかにますます依存している。本サーベイは、エージェント的成果物生成(agentic artifact creation)を検討する。これは、AIシステムが成果物を実質的に構築または改訂し、中間的な観測結果がその後の作業を方向付ける、状態を保持する構築として定義される。機能的には、このプロセスは、成果物の操作可能な表現、構築ポリシー、およびその後のアクションを方向付け得るフィードバックを提供する実行時検証を結びつけるものである。2026年8月20日までに利用可能な259件の研究をレビューした。この定義を満たす230のシステムと、エージェント的成果物構築の29のベンチマークである。6つの成果物ファミリーを比較し、その後、アプリケーション設定と評価実践を別々の次元として分析する。ファミリー全体を通じて、構築の課題はモダリティだけでなく、決定がどの程度緊密に結合されているか、障害が修復可能なうちに可視化されるかどうかにも反映される。分解は局所的な複雑さを低減できる一方で、調整コストと再組み立てコストを増加させる。学習ベースの評価器は、生成器の選好や盲点を共有する場合、独立した証拠をほとんど追加しない可能性がある。我々は、コミットメントと責任を明確に保ち、フィードバックを的を絞った修復に変換し、変更後に影響を受けた状態を再検証するための原則を定式化する。また、成果物、作成者の意図、構築システムが進化するにつれて、一貫性があり説明責任を伴う制御を維持する機会も特定する。厳選された論文リストはhttps://github.com/GeminiLight/awesome-agentic-artifact-creationで入手可能である。
English
Generative models can turn natural-language prompts into images, text, code, and other content, lowering the cost of producing drafts and components. Their practical impact increasingly depends on whether those pieces can become complete, dependable deliverables. This survey examines agentic artifact creation, which we define as stateful construction in which an AI system materially constructs or revises a deliverable and intermediate observations redirect later work. Functionally, the process links an operational representation of the artifact, a construction policy, and runtime verification whose feedback can redirect later actions. We reviewed 259 works available through August 20, 2026: 230 systems meeting this definition and 29 benchmarks of agentic artifact construction. We compare six artifact families, then analyze application settings and evaluation practice as separate dimensions. Across families, construction challenges reflect not only modality but also how tightly decisions are coupled and whether failures become visible while they remain repairable. Decomposition can reduce local complexity while increasing coordination and reassembly costs. Learned judges may add little independent evidence when they share the generator's preferences or blind spots. We formulate principles for keeping commitments and responsibility explicit, turning feedback into targeted repair, and revalidating affected state after change. We also identify opportunities for sustaining coherent, accountable control as artifacts, creator intent, and construction systems evolve. A curated paper list is available at https://github.com/GeminiLight/awesome-agentic-artifact-creation.
PDF563September 1, 2026