ChatPaper.aiChatPaper

MNIST-PRO: MNIST가 AI 에이전트를 위한 부분 관측 가능한 세계로 재탄생하다

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents

August 31, 2026
저자: Vernon Toh, Navonil Majumder, Zhengyuan Liu, Nancy F. Chen, Soujanya Poria
cs.AI

초록

부분 관측 가능한 환경에서 AI 에이전트는 진화하는 지각 상태를 유지하기 위해 능동 감지와 작업 기억을 조정해야 한다. 그러나 기존 벤치마크는 물리적·제어적 복잡성을 도입하여 이러한 지각 상태의 구축 및 해석 능력을 분리하여 평가하는 데 어려움이 있다. 본 연구는 MNIST-PRO를 통해 이 문제를 해결한다. MNIST-PRO는 MNIST 숫자 인식을 회고 제약이 있는 순차적 응시 기반 탐색 과제로 변환하여 에이전트 지각을 분리 평가하는 벤치마크이다. 우리는 원시 시각 이력, 텍스트 상태, 구조화된 미터법 격자 지도, 통합 시각 캔버스의 네 가지 기억 표현에 걸쳐 열 개의 다중 모달 모델을 평가한다. 모델들은 완전 관측 가능성 하에서는 우수한 성능을 보이지만, 부분 관측 가능성은 명확한 성능 격차를 드러낸다. 우리는 세 가지 뚜렷한 병목 현상을 식별한다. 첫째, 에이전트가 단편적인 응시 정보를 통합하는 데 어려움을 겪으면서 지각 상태의 구축과 해석이 과제로 제기된다. 둘째, 에이전트는 전체 시퀀스를 보기 전에 탐색을 중단하는 경우가 많다. 셋째, 모델들은 이후의 상반된 증거에 직면하더라도 초기의 잘못된 신념을 수정하지 못하는 경우가 많다. 이러한 결과는 단순히 시각적 증거를 획득하는 것만으로는 충분하지 않음을 보여준다. 에이전트는 또한 신뢰할 수 있는 지각 상태를 구축하고 갱신할 수 있어야 한다.
English
AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However, existing benchmarks struggle to isolate this perceptual-state construction and interpretation capability because they introduce physical and control complexities. We address this with MNIST-PRO, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints. We evaluate ten multimodal models across four memory representations, including raw visual history, textual states, structured metric grid maps, and a consolidated visual canvas. While models excel under full observability, partial observability exposes a clear performance gap. We identify three distinct bottlenecks. First, perceptual-state construction and interpretation present a challenge, as agents struggle to integrate fragmented glimpses. Second, agents often stop exploring before they see the full sequence. Third, models often fail to revise early, incorrect beliefs even when faced with subsequent contradictory evidence. These results show that simply acquiring visual evidence is not enough. Agents must also be able to build and update a reliable perceptual state.