ChatPaper.aiChatPaper

主动观察者测验

An Exam for Active Observers

July 17, 2026
作者: Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger
cs.AI

摘要

人类视觉是一个闭环过程:注视方向不断被中间假设重新调整,而非依赖单一快照。数十年的心理物理学与认知科学研究表明,这种主动观察对广泛任务至关重要。当前的多模态大语言模型(MLLMs)是否具备主动观察能力,是一个现有视觉-语言基准无法回答的经验性问题。我们提出ActiveVision基准,通过包含3大类17项任务,使MLLMs的主动观察能力得以量化。这些任务强制要求模型进行重复视觉感知,而非仅依赖单次静态描述。前沿MLLMs在ActiveVision上表现崩溃:我们评估的最高分模型——采用最高推理努力层级的GPT-5.5,仅解决10.6%的任务项,且在17项任务中有11项得分为零;而Claude Fable 5尽管在多数推理与编程排行榜上领先,却仅解决3.5%的任务,远低于三位平均得分96.1%的人类参与者。此外,即便模型能自主编写并执行视觉代码,其性能差距仍显著存在:此类代码在真实图像中不可靠,而识别其自身失败恰恰需要模型所缺乏的主动感知能力。综合结果表明,当前MLLMs缺乏稳健的主动视觉观察能力,亟需设计能闭合感知-推理回路的架构与训练目标。
English
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks, and even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.