ChatPaper.aiChatPaper

見たものを描け:マルチモーダルエージェントにおける巧みな視覚的ツール使用のベンチマーキング

Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents

August 26, 2026
著者: Shudong Liu, Dongyang Chen, Enci Zhang, Jinwei Liang, Zheng Ma, Lewei Lu
cs.AI

要旨

評価は静的QAから、モデルが外部ツールを通じて行動するエージェント的な設定へと移行しつつある。我々は、この領域において重要でありながら未開拓の能力、すなわち「器用な視覚ツール使用(dexterous visual tool use)」を特定する。これは、モデルが視覚的証拠からツールのパラメータを推論し、そのパラメータが最終結果を直接決定する、細粒度かつ閉ループのパラメータ化された視覚的行動である。既存のベンチマークはWebナビゲーション、GUI操作、ソフトウェア工学をカバーしているが、視覚的証拠と実行精度のこの結合を対象とすることはほとんどない。我々は、参照ガイド付き視覚再構成を主要な代理タスクとして採用する、器用な視覚ツール使用の制御された事例を評価するベンチマークEASELを提案する。このタスクでは、エージェントが参照画像に一致するようにキャンバスを段階的に描画する。EASELはさらに、領域アノテーション、手書き、経路計画にわたる意味的タスクを含む。さらに、軌跡の教師信号のための44万サンプルの二段階カリキュラムデータセットであるEASEL-Dataと、このデータセットが当該能力に与える効果を調査するためのEASEL-9Bを提供する。25モデルの評価により、現在のマルチモーダルエージェントがEASELで一貫して苦戦することが明らかになる。再構成類似度は低い水準(0.40〜0.54)で頭打ちとなり、軌跡診断は深刻な閉ループ不安定性を露呈する。モデルは典型的には早期に飽和するか、ピーク後に劣化する。意味的タスクは、精密アノテーションと経路計画における明確な能力境界を明らかにする。EASEL-Dataで訓練されたEASEL-9Bは、ベースモデルを相対的な6.3%の向上で上回り、評価された全モデル中第3位となった。
English
Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.