代理式工件創作:系統、評估、原則與機遇

Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities

August 28, 2026
作者: Tianfu Wang, Zhezheng Hao, Xilin Xia, Lixin Liu, Mengkang Hu, Hongzhang Liu, Xi Chen, Ziyan Liu, Xiankun Lin, Weijia Zhang, Nicholas Jing Yuan, Hui Xiong
cs.AI

摘要

生成式模型能將自然語言提示轉化為圖像、文字、程式碼及其他內容,從而降低草稿與元件產製的成本。其實務影響力日益取決於這些片段能否成為完整且可靠的交付成果。本調查檢視「代理式成品創作」(agentic artifact creation),我們將其定義為一種具狀態的建構過程:在此過程中,AI 系統實質地建構或修訂交付成果,且中間觀測結果會引導後續工作。就功能而言,此過程連結了成品的作業表徵、建構策略,以及執行期驗證——其回饋可導引後續動作。我們審閱了截至 2026 年 8 月 20 日可取得的 259 篇文獻:其中 230 個系統符合此定義,另有 29 個代理式成品建構之衡量基準。我們比較了六個成品類別,再分別就應用情境與評測實務兩個面向進行分析。跨類別而言,建構挑戰不僅反映模態差異,也反映決策之間的耦合緊密程度,以及失敗在仍可修復時是否會顯現。分解雖可降低局部複雜度,卻會增加協調與重組的成本。當學習式評判器與生成器共享相同的偏好或盲點時,其所能提供的獨立證據可能相當有限。我們提出若干原則,以確保承諾與責任的明確性、將回饋轉化為標靶式修復,以及在變更後重新驗證受影響的狀態。我們亦指出在成品、創作者意圖與建構系統持續演進之際,維持連貫且可課責之控制的可能性。完整的論文清單請見 https://github.com/GeminiLight/awesome-agentic-artifact-creation。
English
Generative models can turn natural-language prompts into images, text, code, and other content, lowering the cost of producing drafts and components. Their practical impact increasingly depends on whether those pieces can become complete, dependable deliverables. This survey examines agentic artifact creation, which we define as stateful construction in which an AI system materially constructs or revises a deliverable and intermediate observations redirect later work. Functionally, the process links an operational representation of the artifact, a construction policy, and runtime verification whose feedback can redirect later actions. We reviewed 259 works available through August 20, 2026: 230 systems meeting this definition and 29 benchmarks of agentic artifact construction. We compare six artifact families, then analyze application settings and evaluation practice as separate dimensions. Across families, construction challenges reflect not only modality but also how tightly decisions are coupled and whether failures become visible while they remain repairable. Decomposition can reduce local complexity while increasing coordination and reassembly costs. Learned judges may add little independent evidence when they share the generator's preferences or blind spots. We formulate principles for keeping commitments and responsibility explicit, turning feedback into targeted repair, and revalidating affected state after change. We also identify opportunities for sustaining coherent, accountable control as artifacts, creator intent, and construction systems evolve. A curated paper list is available at https://github.com/GeminiLight/awesome-agentic-artifact-creation.
PDF563September 1, 2026