畫你所見:多模態智能體中靈巧視覺工具使用的評測
Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents
August 26, 2026
作者: Shudong Liu, Dongyang Chen, Enci Zhang, Jinwei Liang, Zheng Ma, Lewei Lu
cs.AI
摘要
評估正從靜態問答轉向智能體情境,在這些情境中,模型透過外部工具採取行動。我們在此領域中辨識出一項關鍵卻尚未被充分探索的能力——靈巧的視覺工具使用:一種細粒度、閉環的參數化視覺動作,模型從視覺證據推斷工具參數,而這些參數直接決定最終結果。現有基準涵蓋網頁導航、圖形介面操作與軟體工程,但很少針對視覺證據與執行精度之間的這種耦合。我們提出 EASEL,一個評估靈巧視覺工具使用之受控實例的基準,採用以參考引導的視覺重建作為其主要代理任務:智能體逐步在畫布上繪製,以匹配參考圖像。EASEL 另外包含涵蓋區域標註、手寫與路徑規劃的語義任務。我們進一步提供 EASEL-Data——一個包含 440k 樣本、兩階段課程式的軌跡監督資料集——以及 EASEL-9B,以研究其對此能力的影響。對 25 個模型的評估顯示,當前的多模態智能體在 EASEL 上系統性地表現不佳。重建相似度在低水平(0.40–0.54)出現瓶頸,而軌跡診斷揭露嚴重的閉環不穩定性——模型通常過早飽和或在峰值後退化。語義任務揭示了精確標註與路徑規劃方面鮮明的能力界限。在 EASEL-Data 上訓練的 EASEL-9B 相對基礎模型提升 6.3%,在所有受評模型中排名第三。
English
Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.