能動的観察者のための試験
An Exam for Active Observers
July 17, 2026
著者: Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger
cs.AI
要旨
人間の視覚は閉ループである。すなわち、視線は単一のスナップショットではなく、中間仮説によって継続的に方向転換される。数十年にわたる精神物理学および認知科学の研究により、この能動的観察が幅広いタスクに不可欠であることが論証されてきた。今日のマルチモーダル大規模言語モデル(MLLM)が能動的観察を実行しているかどうかは、現在の視覚言語ベンチマークでは答えられない経験的疑問である。本稿では、MLLMの能動的観察を測定可能にするベンチマーク「ActiveVision」を紹介する。このベンチマークは3カテゴリにわたる17のタスクで構成される。タスクは、単一の静的な記述ではなく、繰り返しの視覚的知覚を強制するように設計されている。最先端のMLLMはActiveVisionで崩壊する。評価した中で最高スコアのモデルであるGPT-5.5(最も露出度の高い推論努力階層)は、項目のわずか10.6%しか解決できず、17タスク中11タスクでスコア0であり、さらにClaude Fable 5は、ほとんどの推論およびコーディングのリーダーボードでトップであるにもかかわらず、わずか3.5%しか解決できず、平均96.1%の3人の人間参加者に大きく劣る。さらに、モデルが自身の視覚コードを記述して実行した場合でも、この差の大部分は残る。そのようなコードは現実的な画像に対して信頼性が低く、その失敗を捉えること自体に、モデルが欠如している能動的知覚が必要だからである。これらの結果は総合して、現在のMLLMには頑健な能動的視覚観察が欠如しており、知覚-推論ループを閉じるアーキテクチャと訓練目的が動機付けられることを示している。
English
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks, and even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.