보이는 대로 그려라: 멀티모달 에이전트의 정교한 시각적 도구 사용 벤치마킹
Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents
August 26, 2026
저자: Shudong Liu, Dongyang Chen, Enci Zhang, Jinwei Liang, Zheng Ma, Lewei Lu
cs.AI
초록
평가는 정적 QA에서 모델이 외부 도구를 통해 작동하는 에이전트적 환경으로 전환되고 있다. 우리는 이 영역 내에서 중요하지만 충분히 탐구되지 않은 능력, 즉 정교한 시각적 도구 사용(dexterous visual tool use)을 식별한다. 이는 모델이 시각적 증거로부터 도구 매개변수를 추론하고, 해당 매개변수가 최종 결과를 직접적으로 좌우하는 세밀하고 폐루프적인 매개변수화된 시각적 행동을 의미한다. 기존 벤치마크는 웹 내비게이션, GUI 조작, 소프트웨어 엔지니어링을 다루지만, 시각적 증거와 실행 정밀도 간의 이러한 결합을 다루는 경우는 드물다. 우리는 EASEL을 제안한다. EASEL은 정교한 시각적 도구 사용의 통제된 사례를 평가하는 벤치마크로, 참조 안내 시각적 재구성(reference-guided visual reconstruction)을 주요 대리 작업으로 채택한다. 에이전트는 기준 이미지와 일치하도록 캔버스를 점진적으로 페인팅한다. EASEL은 또한 영역 주석, 필기, 경로 계획을 포함하는 의미론적 작업을 포함한다. 나아가 궤적 감독을 위한 44만 개 샘플로 구성된 2단계 커리큘럼 데이터셋인 EASEL-Data와 이 능력에 대한 그 효과를 조사하기 위한 EASEL-9B를 제공한다. 25개 모델에 대한 평가 결과, 현재의 멀티모달 에이전트는 EASEL에서 체계적으로 어려움을 겪는 것으로 나타났다. 재구성 유사도는 낮은 수준(0.40-0.54)에서 병목 현상을 보이며, 궤적 진단은 심각한 폐루프 불안정성을 드러낸다. 모델은 일반적으로 조기에 포화되거나 정점 이후 성능이 저하된다. 의미론적 작업은 정밀 주석 및 경로 계획에서 뚜렷한 능력 경계를 드러낸다. EASEL-Data로 훈련된 EASEL-9B는 기준 모델을 상대적으로 6.3% 능가하며 평가된 모든 모델 중 3위를 차지한다.
English
Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.