MNIST-PRO:MNIST作为AI智能体的部分可观测世界回归
MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents
August 31, 2026
作者: Vernon Toh, Navonil Majumder, Zhengyuan Liu, Nancy F. Chen, Soujanya Poria
cs.AI
摘要
在部分可观察环境中,人工智能智能体需要协调主动感知与工作记忆,以维护不断演化的感知状态。然而,现有基准测试因引入了物理和控制复杂性,难以单独评估这种感知状态的构建与解释能力。为此,我们提出了MNIST-PRO,这一基准通过将MNIST数字识别转化为一个带回溯约束的、基于瞥视的序列搜索任务,来隔离智能体感知能力。我们评估了十个多模态模型在四种记忆表示下的表现,包括原始视觉历史、文本状态、结构化度量网格地图以及整合视觉画布。尽管模型在完全可观察条件下表现出色,但部分可观察性暴露了明显的性能差距。我们识别出三个不同的瓶颈。首先,感知状态的构建与解释是一项挑战,因为智能体难以整合碎片化的瞥视信息。其次,智能体往往在看到完整序列之前就停止探索。第三,即使面对后续的矛盾证据,模型也常常无法修正早期形成的错误信念。这些结果表明,仅仅获取视觉证据是不够的。智能体还必须能够构建并更新可靠的感知状态。
English
AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However, existing benchmarks struggle to isolate this perceptual-state construction and interpretation capability because they introduce physical and control complexities. We address this with MNIST-PRO, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints. We evaluate ten multimodal models across four memory representations, including raw visual history, textual states, structured metric grid maps, and a consolidated visual canvas. While models excel under full observability, partial observability exposes a clear performance gap. We identify three distinct bottlenecks. First, perceptual-state construction and interpretation present a challenge, as agents struggle to integrate fragmented glimpses. Second, agents often stop exploring before they see the full sequence. Third, models often fail to revise early, incorrect beliefs even when faced with subsequent contradictory evidence. These results show that simply acquiring visual evidence is not enough. Agents must also be able to build and update a reliable perceptual state.