에이전트 기반 산출물 생성: 시스템, 평가, 원칙 및 기회
Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
August 28, 2026
저자: Tianfu Wang, Zhezheng Hao, Xilin Xia, Lixin Liu, Mengkang Hu, Hongzhang Liu, Xi Chen, Ziyan Liu, Xiankun Lin, Weijia Zhang, Nicholas Jing Yuan, Hui Xiong
cs.AI
초록
생성 모델은 자연어 프롬프트를 이미지, 텍스트, 코드 및 기타 콘텐츠로 변환하여 초안과 구성 요소를 생산하는 비용을 낮춘다. 이러한 모델의 실질적 영향력은 점점 더 그러한 조각들이 완전하고 신뢰할 수 있는 인도물로 전환될 수 있는지에 달려 있다. 본 조사(survey)는 에이전트 기반 산출물 생성(agentic artifact creation)을 검토하며, 이를 AI 시스템이 인도물을 실질적으로 생성하거나 수정하고 중간 관찰이 이후 작업을 방향 전환시키는 상태 유지형 구축(stateful construction)으로 정의한다. 기능적으로 이 과정은 산출물의 운영적 표현(operational representation), 구축 정책(construction policy), 그리고 피드백을 통해 이후 행동을 방향 전환할 수 있는 런타임 검증(runtime verification)을 연결한다. 우리는 2026년 8월 20일까지 이용 가능한 259편의 연구를 검토하였다: 이 정의를 충족하는 230개 시스템과 에이전트 기반 산출물 구축을 위한 29개 벤치마크이다. 우리는 여섯 가지 산출물 계열(artifact family)을 비교한 다음, 애플리케이션 설정과 평가 관행을 별도의 차원으로 분석한다. 계열 전반에 걸쳐 구축 과제는 양식(modality)뿐만 아니라 결정이 얼마나 긴밀하게 결합되어 있는지, 그리고 실패가 여전히 수리 가능한 동안 가시화되는지 여부에 따라 달라진다. 분해(decomposition)는 지역적 복잡성을 줄일 수 있지만 조정 및 재조립 비용은 증가시킨다. 학습된 평가기(learned judge)는 생성기의 선호도나 맹점을 공유할 때 독립적 증거를 거의 추가하지 못할 수 있다. 우리는 약속과 책임을 명시적으로 유지하고, 피드백을 목표 지향적 수리로 전환하며, 변경 후 영향받은 상태를 재검증하기 위한 원칙을 정립한다. 또한 산출물, 창작자 의도, 구축 시스템이 진화함에 따라 지속 가능하고 책임 있는 통제를 유지하기 위한 기회를 식별한다. 선별된 논문 목록은 https://github.com/GeminiLight/awesome-agentic-artifact-creation 에서 확인할 수 있다.
English
Generative models can turn natural-language prompts into images, text, code, and other content, lowering the cost of producing drafts and components. Their practical impact increasingly depends on whether those pieces can become complete, dependable deliverables. This survey examines agentic artifact creation, which we define as stateful construction in which an AI system materially constructs or revises a deliverable and intermediate observations redirect later work. Functionally, the process links an operational representation of the artifact, a construction policy, and runtime verification whose feedback can redirect later actions. We reviewed 259 works available through August 20, 2026: 230 systems meeting this definition and 29 benchmarks of agentic artifact construction. We compare six artifact families, then analyze application settings and evaluation practice as separate dimensions. Across families, construction challenges reflect not only modality but also how tightly decisions are coupled and whether failures become visible while they remain repairable. Decomposition can reduce local complexity while increasing coordination and reassembly costs. Learned judges may add little independent evidence when they share the generator's preferences or blind spots. We formulate principles for keeping commitments and responsibility explicit, turning feedback into targeted repair, and revalidating affected state after change. We also identify opportunities for sustaining coherent, accountable control as artifacts, creator intent, and construction systems evolve. A curated paper list is available at https://github.com/GeminiLight/awesome-agentic-artifact-creation.