ChatPaper.aiChatPaper

MNIST-PRO: 部分観測可能なAIエージェント向け環境としてMNISTが再登場

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents

August 31, 2026
著者: Vernon Toh, Navonil Majumder, Zhengyuan Liu, Nancy F. Chen, Soujanya Poria
cs.AI

要旨

部分観測可能環境におけるAIエージェントは、進化する知覚状態を維持するために、能動的センシングを作業記憶と協調させる必要がある。しかしながら、既存のベンチマークは、物理的・制御的な複雑性を導入するため、この知覚状態の構築および解釈能力を切り離して評価することが困難である。我々はこの問題に対処するため、MNIST-PROを提案する。これは、MNIST数字認識を、ルックバック制約を伴う逐次的なグリンプスベース探索タスクへと変換することで、エージェント的知覚を切り離して評価するベンチマークである。我々は、生の視覚履歴、テキストベースの状態表現、構造化されたメトリックグリッドマップ、統合された視覚キャンバスという4つの記憶表現にわたって、10のマルチモーダルモデルを評価する。モデルは完全観測下では優れた性能を示す一方、部分観測下では明確な性能ギャップが顕在化する。我々は3つの明確なボトルネックを特定する。第一に、断片的なグリンプスを統合することにエージェントが困難を示すため、知覚状態の構築と解釈が課題となる。第二に、エージェントは完全なシーケンスを確認する前に探索を終了してしまうことが多い。第三に、モデルはその後の矛盾する証拠に直面しても、初期の誤った信念を修正できないことが多い。これらの結果は、単に視覚的証拠を取得するだけでは不十分であることを示している。エージェントは信頼性の高い知覚状態を構築し、更新する能力も備えていなければならないのである。
English
AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However, existing benchmarks struggle to isolate this perceptual-state construction and interpretation capability because they introduce physical and control complexities. We address this with MNIST-PRO, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints. We evaluate ten multimodal models across four memory representations, including raw visual history, textual states, structured metric grid maps, and a consolidated visual canvas. While models excel under full observability, partial observability exposes a clear performance gap. We identify three distinct bottlenecks. First, perceptual-state construction and interpretation present a challenge, as agents struggle to integrate fragmented glimpses. Second, agents often stop exploring before they see the full sequence. Third, models often fail to revise early, incorrect beliefs even when faced with subsequent contradictory evidence. These results show that simply acquiring visual evidence is not enough. Agents must also be able to build and update a reliable perceptual state.