ChatPaper.aiChatPaper

See2Think: マルチモーダルモデルは中間視覚状態を本当に利用しているのか?

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

July 29, 2026
著者: Siyu Yan, Zhuoran Yan, Haiying Xu, Panhao Zhou, Jingyu Chen, Chenhao Ji, Shuo Cao, Yongheng Zhang, Haoze Liu, Siyu Zhang, Xiwen Gu, Yihao Liu, Alex Jinpeng Wang
cs.AI

要旨

マルチモーダル大規模言語モデルは、推論中にスケッチ、注釈、ツール、中間画像をますます活用しているが、これらの視覚的状態に本当に依存しているかは依然として不明である。既存のベンチマークは、カバレッジが狭いタスクコレクションやテキストのみで部分的に解けるサンプルに限定されているだけでなく、中間的な視覚的状態がどのように生成、描画、利用されるかを診断することなく最終回答のみを重視する評価にも制約されている。我々は、See2ThinkBenchとVisual Action-of-Thought(VAoT)からなる統合評価フレームワークであるSee2Thinkを提案する。See2ThinkBenchには、2D構造化推論、3Dシーン推論、実世界推論にわたる12のタスクカテゴリに及ぶ1,200問のオープンエンドで視覚依存的問題が含まれる。VAoTは、4つの制御された推論設定下で、テキストによる思考、視覚的行動、描画された状態、その後の推論を記録する。代表的ないくつかのプロプライエタリおよびオープンソースのマルチモーダルモデルを評価した結果、視覚的推論はモデルと環境に強く依存し、タスク間で一貫して優位な単一の設定は存在しないことが判明した。プロセス分析により、モデルは通常関連する視覚的操作を選択する一方で、忠実な描画が最も明確なボトルネックであり、高いフィードバック取り込み率が必ずしも精度向上につながらないことがさらに示された。タスク関連の破損フィードバック下では、モデルは視覚的状態への行動的依存を示し、制御された介入において精度が10パーセントポイント以上低下した。
English
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.