ChatPaper.aiChatPaper

MNIST-PRO:MNIST以部分可觀測世界之姿回歸,面向AI智能體

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents

August 31, 2026
作者: Vernon Toh, Navonil Majumder, Zhengyuan Liu, Nancy F. Chen, Soujanya Poria
cs.AI

摘要

在部分可觀測環境中,AI代理需要協調主動感知與工作記憶,以維持持續演進的感知狀態。然而,現有基準測試因引入了物理與控制上的複雜度,難以分離出這種感知狀態建構與詮釋的能力。我們提出 MNIST-PRO 來解決此問題;此基準測試將 MNIST 數字辨識轉化為一個具回溯限制、依序進行瞥視的搜尋任務,藉此分離出代理式感知能力。我們在四種記憶表徵下評估了十個多模態模型,包括原始視覺歷史、文字狀態、結構化度量網格地圖,以及整合式視覺畫布。儘管模型在完全可觀測下表現優異,部分可觀測性卻暴露出明顯的效能落差。我們辨識出三個不同的瓶頸。首先,感知狀態的建構與詮釋構成一項挑戰,因為代理難以整合碎片化的瞥視。其次,代理往往在看完完整序列之前就停止探索。第三,即使面對後續相互矛盾的證據,模型仍常無法修正早期錯誤的信念。這些結果顯示,僅僅獲取視覺證據並不足夠;代理還必須能夠建立並更新可靠的感知狀態。
English
AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However, existing benchmarks struggle to isolate this perceptual-state construction and interpretation capability because they introduce physical and control complexities. We address this with MNIST-PRO, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints. We evaluate ten multimodal models across four memory representations, including raw visual history, textual states, structured metric grid maps, and a consolidated visual canvas. While models excel under full observability, partial observability exposes a clear performance gap. We identify three distinct bottlenecks. First, perceptual-state construction and interpretation present a challenge, as agents struggle to integrate fragmented glimpses. Second, agents often stop exploring before they see the full sequence. Third, models often fail to revise early, incorrect beliefs even when faced with subsequent contradictory evidence. These results show that simply acquiring visual evidence is not enough. Agents must also be able to build and update a reliable perceptual state.