ChatPaper.aiChatPaper

心智世界建模

Mental World Modeling

July 29, 2026
作者: Hao Fei, Yiran Zhao
cs.AI

摘要

世界模型为规划和行动提供了预测性基础,然而现有公式仅仅回答了一个物理问题:事物是什么/在哪里,以及它将如何演化。然而,人类行为是由隐藏的心智状态(一个人相信什么、想要什么、意图做什么、感受什么,以及认为什么在社会层面是可被允许的)所驱动的,因此,一个跟踪物理场景但不跟踪每个智能体对其所知和所信的模型,会为看似正确的场景预测出错误的行动。我们提出了心智世界建模(MWM),这是一个通用的理论框架,将心智变量作为世界模型的核心组成部分,而非事后的解释理由:MWM 维持一个耦合的物理-心智世界状态,渲染目标特定的部分观察,并模拟候选行动如何联合更新这两个组成部分。我们在 MENTIS 中实例化了该框架,这是一个免训练且完全可检查的基线,它将过程分解为状态解析、目标观察生成、行动分解、耦合的物理与心智转移,以及分支级价值评估。在一个手工构建、质量控制的情境化决策场景数据集上(涵盖文本、图像和有声视频故事),使用 8 个基于现代大语言模型的世界模型进行的实验表明,显式建模心智状态对于预测人类决策至关重要。更深入的分析进一步揭示了当前心智世界建模的瓶颈。我们期望 MWM 成为世界建模的下一个阶段,从模拟物理场景转向模拟在其中行动的智能体的心智。
English
World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, and considers socially permissible), so a model that tracks the physical scene but not what each agent knows and believes about it predicts the wrong action for the right-looking scene. We formulate Mental World Modeling (MWM), a generic theoretical framework that makes mental variables core components of a world model rather than posthoc rationales: MWM aintains a coupled physical-mental world state, renders a target-specific partial observation, and simulates how candidate actions jointly update both components. We instantiate the framework in MENTIS, a training-free and fully inspectable baseline that decomposes the process into state parsing, target-observation generation, action decomposition, coupled physical and mental transition, and branch-level value evaluation. On a manually constructed, quality-controlled dataset of situated decision scenarios spanning text, image, and sounding-video stories, experiments with 8 modern LLM-based world models demonstrate that explicitly modeling the mental state is essential for predicting human decisions. Deeper analyses further expose the bottlenecks of current mental world modeling. We expect MWM as a next stage of world modeling, from simulating physical scenes to simulating the minds that act in them.