ChatPaper.aiChatPaper

See2Think: 멀티모달 모델은 실제로 중간 시각 상태를 사용하는가?

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

July 29, 2026
저자: Siyu Yan, Zhuoran Yan, Haiying Xu, Panhao Zhou, Jingyu Chen, Chenhao Ji, Shuo Cao, Yongheng Zhang, Haoze Liu, Siyu Zhang, Xiwen Gu, Yihao Liu, Alex Jinpeng Wang
cs.AI

초록

멀티모달 대규모 언어 모델은 추론 과정에서 스케치, 주석, 도구, 중간 이미지를 점점 더 많이 활용하지만, 이러한 시각적 상태에 실제로 의존하는지 여부는 여전히 불분명하다. 기존 벤치마크는 좁은 범위를 다루는 작업 컬렉션이나 부분적으로 텍스트만으로 해결 가능한 샘플로 인해 제한적이며, 중간 시각적 상태가 어떻게 생성, 렌더링, 활용되는지를 진단하지 않고 최종 답변만을 평가하는 데 초점을 맞춘다는 한계가 있다. 우리는 See2ThinkBench와 Visual Action-of-Thought(VAoT)로 구성된 통합 평가 프레임워크인 See2Think를 제안한다. See2ThinkBench는 2D 구조적 추론, 3D 장면 추론, 실세계 추론을 포괄하는 12개 작업 범주에 걸친 1,200개의 개방형 시각 의존 문제를 포함한다. VAoT는 네 가지 통제된 추론 설정 하에서 텍스트 기반 사고, 시각적 행동, 렌더링된 상태, 후속 추론을 기록한다. 대표적인 상용 및 오픈소스 멀티모달 모델을 평가한 결과, 시각적 추론은 모델과 환경에 강하게 의존하며 작업 전반에서 단일 설정이 일관되게 우세하지 않음을 발견했다. 과정 분석은 모델이 일반적으로 관련 시각적 작업을 선택하지만 충실한 렌더링이 가장 명확한 병목 현상으로 남아 있으며, 높은 피드백 수용률이 반드시 정확도 향상으로 이어지지는 않음을 보여준다. 작업 관련 손상된 피드백이 주어진 상황에서 모델은 시각적 상태에 대한 행동적 의존성을 나타내며, 통제된 개입 하에서 정확도가 10퍼센트포인트 이상 하락한다.
English
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.