ChatPaper.aiChatPaper

主動觀察者測驗

An Exam for Active Observers

July 17, 2026
作者: Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger
cs.AI

摘要

人類視覺是一個閉環:目光會基於中間假設不斷重新定向,而非一次性快照。數十年的心理物理學與認知科學研究指出,這種主動觀察對廣泛任務至關重要。當今多模態大語言模型(MLLM)是否具備主動觀察能力,是一個現有視覺語言基準測試無法解答的實證問題。我們提出ActiveVision,這是一個能衡量MLLM主動觀察能力的基準測試,包含三大類別共17項任務。任務設計旨在強迫模型進行重複性的視覺感知,而非依賴單一靜態描述。尖端MLLM在ActiveVision上表現崩潰:我們評估中得分最高的模型GPT-5.5(暴露於最高推理努力層級),僅能解決10.6%的項目,並在17項任務中有11項得零分;即便是在多數推理與程式碼排行榜上位居首位的Claude Fable 5,也只解決了3.5%,遠落後於三位平均得分96.1%的人類參與者。此外,即使模型撰寫並執行自己的視覺程式碼,大部分差距依然存在:這類程式碼在真實影像上不可靠,且偵測其失誤本身就需要模型所缺乏的主動感知能力。綜合這些結果顯示,當前MLLM缺乏穩健的主動視覺觀察能力,這促使我們需要設計能閉合感知-推理迴路的架構與訓練目標。
English
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks, and even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.