See2Think:多模态模型真的使用中间视觉状态吗?
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
July 29, 2026
作者: Siyu Yan, Zhuoran Yan, Haiying Xu, Panhao Zhou, Jingyu Chen, Chenhao Ji, Shuo Cao, Yongheng Zhang, Haoze Liu, Siyu Zhang, Xiwen Gu, Yihao Liu, Alex Jinpeng Wang
cs.AI
摘要
多模态大语言模型在推理过程中越来越多地使用草图、标注、工具和中间图像,但目前尚不清楚它们是否真正依赖这些视觉状态。现有基准测试既受限于任务集合覆盖范围窄或部分可通过文本求解的样本,也受限于评估方式仅关注最终答案而未诊断中间视觉状态如何生成、渲染和使用。我们提出See2Think,一个统一的评估框架,包含See2ThinkBench和视觉思维动作(Visual Action-of-Thought,VAoT)。See2ThinkBench包含1,200个开放式、依赖视觉的问题,涵盖12个任务类别,涉及2D结构化、3D场景和真实世界推理。VAoT在四种受控推理设置下记录文本思维、视觉动作、渲染状态及后续推理。通过对具有代表性的专有和开源多模态模型进行评估,我们发现视觉推理强烈依赖于模型和环境,没有任何单一设置能在各任务中持续占优。过程分析进一步表明,模型通常会选择相关的视觉操作,而忠实渲染仍是最明显的瓶颈,且高反馈采纳率并不必然转化为准确率的提升。在任务相关的扰动反馈条件下,模型表现出对视觉状态的行为依赖性,在受控干预中准确率下降超过10个百分点。
English
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.