ChatPaper.aiChatPaper

에이전틱 시각 생성: 생성 모델에서 에이전틱 제어로

Agentic Visual Generation: From Generative Models to Agentic Control

September 6, 2026
저자: Yinming Huang, Shuyuan Tu, Xi Yan, Jiahao Zhan, Zihan Yang, Zhen Xing, Hui Zhang, Tiehua Zhang, Yu-Gang Jiang, Zuxuan Wu
cs.AI

초록

시각 생성은 단일 호출을 통해 사용되는 생성 모델에서 계획하고, 도구를 선택하고, 중간 합성 출력을 검사하고, 실패를 수정하고, 이전 경험을 재사용할 수 있는 에이전트적 제어 과정으로 진화하고 있다. 대부분의 기존 시스템에서 컨트롤러는 LLM 또는 VLM이며, 시각 생성 모델은 도구 또는 실행기 역할을 한다. 그러나 기존 연구는 생성 시스템이 언제 에이전트적이 되는지를 판단하기 위한 일관된 기준을 결여하고 있다. 계획 깊이, 도구 사용, 다중 역할 협업, 강화학습은 흔히 에이전트성의 증거로 취급되지만, 그중 어느 것도 컨트롤러가 어떤 생성 결정을 내릴 수 있는지를 반드시 결정하지는 않는다. 우리는 생성 과정에서 컨트롤러가 직접 제어할 수 있는 대상에 따라 이 분야를 체계화한다. L1 조건화 제어에서 컨트롤러는 미리 정해진 생성기에 대한 입력을 준비하지만, 어떤 시각적 연산이 수행되는지는 제어하지 않는다. L2 실행 제어에서는 실제 생성, 편집, 렌더링 또는 기타 콘텐츠 수정 연산을 선택하고 호출한다. L3 결과 적응형 제어에서는 중간 결과를 관찰하고, 그 관찰을 사용하여 현재 작업 내의 후속 연산을 변경한다. L4 경험 적응형 제어에서는 완료된 작업의 경험을 유지하고, 그 경험을 사용하여 미래 작업에 대한 결정을 변경한다. L0 고정 지원은 생성 수준 결정을 내리는 배치된 컨트롤러 없이 생성기, 편집기, 평가기, 보상 모델, 벤치마크, 고정 파이프라인을 별도로 나타낸다. 이 수준들은 모델 크기, 시스템 복잡도, 출력 품질, 도구나 역할 수, 또는 학습 방법이 아니라 점진적으로 더 넓어지는 의사결정 범위를 설명한다. 이 프레임워크를 이미지, 비디오, 편집, 3D, 월드, 슬라이드, 사용자 인터페이스 생성 전반에 적용하면 컨트롤러 역량이 어떻게 진화해 왔는지와 그 메커니즘이 수준별로 어떻게 분포하는지를 드러낸다.
English
Visual generation is evolving from generative models used through a single invocation into agentic control processes that can plan, select tools, inspect intermediate synthesized outputs, revise failures, and reuse prior experience. In most existing systems, the controller is an LLM or VLM, while visual generation models serve as tools or executors. However, existing work lacks a consistent criterion for determining when a generation system becomes agentic. Planning depth, tool use, multi-role collaboration, and reinforcement learning are often treated as evidence of agenticity, even though none of them necessarily determines which generation decisions the controller can make. We organize the field according to what the controller can directly control in the generation process. At L1 Conditioning Control, the controller prepares the input to a predetermined generator but does not control which visual operation is executed. At L2 Execution Control, it selects and invokes actual generation, editing, rendering, or other content-modifying operations. At L3 Outcome-Adaptive Control, it observes an intermediate outcome and uses that observation to change a subsequent operation within the current task. At L4 Experience-Adaptive Control, it retains experience from completed tasks and uses that experience to change decisions on future tasks. L0 Fixed Support separately denotes generators, editors, evaluators, reward models, benchmarks, and fixed pipelines without a deployed controller that makes generation-level decisions. These levels describe a progressively broader decision-making scope rather than model size, system complexity, output quality, tool or role count, or training method. Applying this framework across image, video, editing, 3D, world, slide, and user-interface generation reveals how controller capabilities have evolved and how their mechanisms are distributed across levels.