智能体工件创建:系统、评估、原则与机遇

Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities

August 28, 2026
作者: Tianfu Wang, Zhezheng Hao, Xilin Xia, Lixin Liu, Mengkang Hu, Hongzhang Liu, Xi Chen, Ziyan Liu, Xiankun Lin, Weijia Zhang, Nicholas Jing Yuan, Hui Xiong
cs.AI

摘要

生成模型可以将自然语言提示转化为图像、文本、代码及其他内容,从而降低草稿与组件的生成成本。其实际影响力日益取决于这些产物能否成为完整、可靠的交付物。本综述审视了智能体工件创建(agentic artifact creation),我们将其定义为一种有状态构建过程:由AI系统实质性构建或修订交付物,且中间观察结果会重定向后续工作。在功能层面,该过程将工件的操作表示、构建策略与运行时验证相连接,后者的反馈可重定向后续操作。我们综述了截至2026年8月20日可获取的259篇文献:其中230个系统符合这一定义,另有29个用于智能体工件构建的基准。我们比较了六个工件族,随后将应用场景与评测实践作为独立维度进行分析。跨工件族而言,构建挑战不仅反映模态差异,还取决于决策之间的耦合紧密程度,以及失败是否在仍可修复时便被察觉。分解可以降低局部复杂性,但会增加协调与重组成本。当学习型评判器与生成器共享偏好或盲区时,其所能提供的独立证据可能极为有限。我们提出了若干原则:保持承诺与责任明确、将反馈转化为有针对性的修复、并在变更后重新验证受影响的状态。同时,我们也识别出随着工件、创作者意图与构建系统的演化,维持连贯且可追责控制的机会。精选论文列表见 https://github.com/GeminiLight/awesome-agentic-artifact-creation。
English
Generative models can turn natural-language prompts into images, text, code, and other content, lowering the cost of producing drafts and components. Their practical impact increasingly depends on whether those pieces can become complete, dependable deliverables. This survey examines agentic artifact creation, which we define as stateful construction in which an AI system materially constructs or revises a deliverable and intermediate observations redirect later work. Functionally, the process links an operational representation of the artifact, a construction policy, and runtime verification whose feedback can redirect later actions. We reviewed 259 works available through August 20, 2026: 230 systems meeting this definition and 29 benchmarks of agentic artifact construction. We compare six artifact families, then analyze application settings and evaluation practice as separate dimensions. Across families, construction challenges reflect not only modality but also how tightly decisions are coupled and whether failures become visible while they remain repairable. Decomposition can reduce local complexity while increasing coordination and reassembly costs. Learned judges may add little independent evidence when they share the generator's preferences or blind spots. We formulate principles for keeping commitments and responsibility explicit, turning feedback into targeted repair, and revalidating affected state after change. We also identify opportunities for sustaining coherent, accountable control as artifacts, creator intent, and construction systems evolve. A curated paper list is available at https://github.com/GeminiLight/awesome-agentic-artifact-creation.
PDF563September 1, 2026