ChatPaper.aiChatPaper

능동적 관찰자를 위한 시험

An Exam for Active Observers

July 17, 2026
저자: Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger
cs.AI

초록

인간의 시각은 폐쇄 루프(closed loop)로 작동한다: 시선은 단일 스냅샷이 아닌 중간 가설들에 의해 지속적으로 재조정된다. 수십 년간의 심리물리학 및 인지과학 연구는 이러한 능동적 관찰이 다양한 과제 수행에 필수적이라는 점을 주장해 왔다. 오늘날의 다중 모달 대규모 언어 모델(MLLM)이 능동적 관찰을 수행하는지 여부는 현재의 시각-언어 벤치마크로는 답할 수 없는 경험적 질문이다. 본 연구에서는 MLLM의 능동적 관찰을 측정 가능하게 만드는 벤치마크인 ActiveVision을 소개한다. 이 벤치마크는 3개 범주에 걸쳐 17개의 과제로 구성된다. 과제들은 단일 정적 설명이 아닌 반복적인 시각적 지각을 강제하도록 설계되었다. 최첨단 MLLM들은 ActiveVision에서 성능이 급락한다: 우리가 평가한 최고 점수 모델인 GPT-5.5(최고 수준의 추론 노력 단계)는 항목의 10.6%만을 해결했으며, 17개 과제 중 11개에서 0점을 기록했다. 또한 대부분의 추론 및 코딩 리더보드에서 1위를 차지한 Claude Fable 5조차 단 3.5%만을 해결하여, 평균 96.1%를 기록한 세 명의 인간 참가자에 크게 뒤처졌다. 더 나아가, 모델이 자체적으로 시각 코드를 작성하고 실행하더라도 이러한 격차의 상당 부분이 지속된다: 이러한 코드는 실제 이미지에서 신뢰할 수 없으며, 실패를 포착하는 것 자체가 모델이 결여한 능동적 지각을 필요로 한다. 종합하면, 이러한 결과는 현재의 MLLM이 견고한 능동적 시각 관찰을 결여하고 있음을 시사하며, 지각-추론 루프를 폐쇄하는 아키텍처와 훈련 목표의 필요성을 제기한다.
English
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks, and even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.