CanvasAgent:透過視覺工具編排實現複雜圖像創作與編輯
CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
July 6, 2026
作者: Hairui Zhu, Yiying Yang, Tengjin Weng, Ziyu Lu, Xiao Yao, Xiaoyang Ye, Lin Ma, Wenhao Jiang
cs.AI
摘要
複雜的影像創作與編輯往往需要的不僅是單一的生成或編輯模型。使用者需求可能包含合成影像、定位物件、分割區域、編輯選取內容、組合中介素材、讀取文字,以及強化最終成果。這類任務將多模態代理從增強感知的推理,轉變為以操作為核心的視覺創作,其中的工具必須主動改變視覺狀態,而非單純進行檢視。然而,現有的多模態工具使用代理大多針對感知、搜尋或特定領域的編輯進行最佳化,缺乏大規模可執行影像創作軌跡的監督訊號。本文提出 CanvasCraft,這是一個大規模的多模態工具使用資料集,用於複雜影像創作與編輯,以及 CanvasAgent,這是一個工具增強的的多模態代理,能夠學習透過多輪互動來協調異質視覺工具。CanvasCraft 包含 14 萬條完整標註的可執行軌跡,以及 1 萬個強化學習任務規格。CanvasAgent 首先經由 SFT 訓練,學習可執行的推理-行動軌跡,接著使用 GRPO 進行最佳化,採用結合結果層級與過程層級訊號的混合獎勵。在推論過程中,CanvasAgent 會檢查中間結果、追蹤視覺資產,並根據不斷演變的視覺狀態調整工具決策。實驗評估了最終影像品質與軌跡行為,證明了 CanvasAgent 與所提出資料集在複雜多工具影像創作流程中的有效性。
English
Complex image creation and editing often require more than a single generation or editing model. A user request may involve synthesizing images, localizing objects, segmenting regions, editing selected content, compositing intermediate assets, reading text, and enhancing the final result. Such tasks shift multimodal agents from perception-augmented reasoning to manipulation-centered visual creation, where tools must actively transform visual states rather than merely inspect them. However, existing multimodal tool-use agents are mostly optimized for perception, search, or domain-specific editing, and lack large-scale supervision for executable image-creation trajectories. In this paper, we introduce CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and CanvasAgent, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interaction. CanvasCraft contains 140K fully annotated executable trajectories and 10K
RL task specifications. CanvasAgent is first trained with SFT to learn executable reasoning-action trajectories, and is then optimized with GRPO using a hybrid reward that combines outcome- and process-level signals. During rollout, CanvasAgent inspects intermediate results, tracks visual assets, and adapts tool decisions to the evolving visual state. Experiments evaluate both final image quality and trajectory behavior, demonstrating the effectiveness of CanvasAgent and the proposed dataset for complex multi-tool image creation workflows.