CanvasAgent: 시각적 도구 오케스트레이션을 통한 복잡한 이미지 생성 및 편집
CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
July 6, 2026
저자: Hairui Zhu, Yiying Yang, Tengjin Weng, Ziyu Lu, Xiao Yao, Xiaoyang Ye, Lin Ma, Wenhao Jiang
cs.AI
초록
복잡한 이미지 생성 및 편집은 단일 생성 또는 편집 모델만으로는 충분하지 않은 경우가 많습니다. 사용자 요청은 이미지 합성, 객체 위치 파악, 영역 분할, 선택 콘텐츠 편집, 중간 자산 합성, 텍스트 읽기, 최종 결과물 향상 등을 포함할 수 있습니다. 이러한 작업은 다중 모달 에이전트를 지각 기반 추론에서 조작 중심의 시각적 창작으로 전환시키며, 여기서 도구는 단순히 시각적 상태를 검사하는 대신 적극적으로 변환해야 합니다. 그러나 기존의 다중 모달 도구 사용 에이전트는 대부분 지각, 검색 또는 특정 도메인 편집에 최적화되어 있으며, 실행 가능한 이미지 생성 궤적에 대한 대규모 감독이 부족합니다. 본 논문에서는 복잡한 이미지 생성 및 편집을 위한 대규모 다중 모달 도구 사용 데이터셋인 CanvasCraft와, 다중 턴 상호작용을 통해 이기종 시각적 도구를 조율하는 방법을 학습하는 도구 증강 다중 모달 에이전트인 CanvasAgent를 소개합니다. CanvasCraft는 140K개의 완전 주석 처리된 실행 가능한 궤적과 10K개의 RL 작업 사양을 포함합니다. CanvasAgent는 먼저 SFT로 학습되어 실행 가능한 추론-행동 궤적을 학습한 후, 결과 수준 및 과정 수준 신호를 결합한 하이브리드 보상을 사용하는 GRPO로 최적화됩니다. 롤아웃 과정에서 CanvasAgent는 중간 결과를 검사하고, 시각적 자산을 추적하며, 진화하는 시각적 상태에 맞게 도구 결정을 조정합니다. 실험은 최종 이미지 품질과 궤적 행동을 모두 평가하여, 복잡한 다중 도구 이미지 생성 워크플로우에서 CanvasAgent와 제안된 데이터셋의 효과성을 입증합니다.
English
Complex image creation and editing often require more than a single generation or editing model. A user request may involve synthesizing images, localizing objects, segmenting regions, editing selected content, compositing intermediate assets, reading text, and enhancing the final result. Such tasks shift multimodal agents from perception-augmented reasoning to manipulation-centered visual creation, where tools must actively transform visual states rather than merely inspect them. However, existing multimodal tool-use agents are mostly optimized for perception, search, or domain-specific editing, and lack large-scale supervision for executable image-creation trajectories. In this paper, we introduce CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and CanvasAgent, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interaction. CanvasCraft contains 140K fully annotated executable trajectories and 10K
RL task specifications. CanvasAgent is first trained with SFT to learn executable reasoning-action trajectories, and is then optimized with GRPO using a hybrid reward that combines outcome- and process-level signals. During rollout, CanvasAgent inspects intermediate results, tracks visual assets, and adapts tool decisions to the evolving visual state. Experiments evaluate both final image quality and trajectory behavior, demonstrating the effectiveness of CanvasAgent and the proposed dataset for complex multi-tool image creation workflows.