See2Think:多模態模型是否真的使用中間視覺狀態?
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
July 29, 2026
作者: Siyu Yan, Zhuoran Yan, Haiying Xu, Panhao Zhou, Jingyu Chen, Chenhao Ji, Shuo Cao, Yongheng Zhang, Haoze Liu, Siyu Zhang, Xiwen Gu, Yihao Liu, Alex Jinpeng Wang
cs.AI
摘要
多模態大型語言模型在推理過程中日益頻繁地使用草圖、註解、工具及中間圖像,但目前仍不清楚它們是否真正依賴這些視覺狀態。現有基準的侷限性體現在兩個方面:任務集合覆蓋範圍狹窄或包含部分可透過文字解決的樣本,以及評估側重於最終答案而未診斷中間視覺狀態如何被生成、渲染及使用。我們提出 See2Think,一個統一的評估框架,包含 See2ThinkBench 與視覺思維行動(Visual Action-of-Thought, VAoT)。See2ThinkBench 包含 1,200 個開放式、依賴視覺的問題,涵蓋 12 個任務類別,橫跨二維結構、三維場景與真實世界推理。VAoT 則在四種受控推理設定下記錄文字思維、視覺行動、渲染狀態及後續推理。透過評估具代表性的專有與開源多模態模型,我們發現視覺推理高度依賴模型與環境,沒有任何單一設定能在各任務中持續佔據主導地位。過程分析進一步表明,模型通常會選擇相關的視覺操作,而忠實渲染仍是最明顯的瓶頸,且高回饋吸收率並不一定轉化為準確度的提升。在任務相關的受損回饋條件下,模型表現出對視覺狀態的行為依賴,在控制性干預下準確度下降超過 10 個百分點。
English
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.