心智世界建模
Mental World Modeling
July 29, 2026
作者: Hao Fei, Yiran Zhao
cs.AI
摘要
世界模型為規劃與行動提供了可預測的基礎,但現有表述僅回答物理性問題:它是什麼/在哪裡,以及它將如何演變。然而,人類行為是由隱藏的心智狀態所驅動(一個人相信什麼、想要什麼、意圖做什麼、感受如何,以及認為什麼在社會上是被允許的),因此,一個僅追蹤物理場景、卻不追蹤每個智能體對場景所知與所信的模型,會在看似正確的場景中預測出錯誤的行動。我們提出了心智世界建模(Mental World Modeling, MWM),這是一個通用的理論框架,將心智變量作為世界模型的核心組成部分,而非事後解釋:MWM 維持一個耦合的物理-心智世界狀態,生成針對特定目標的部分觀測,並模擬候選行動如何共同更新這兩個組成部分。我們以 MENTIS 實例化該框架,這是一個免訓練且完全可檢視的基線模型,將過程分解為狀態解析、目標觀測生成、行動分解、物理與心智的耦合轉移,以及分支層級的價值評估。在一個人工建構、品質受控的情境化決策場景資料集上,該資料集涵蓋文字、圖像與有聲影片故事,使用 8 個現代基於大型語言模型(LLM)的世界模型所進行的實驗表明,明確建模心智狀態對於預測人類決策至關重要。更深層的分析進一步揭示了當前心智世界建模的瓶頸。我們期望 MWM 能成為世界建模的下一階段,從模擬物理場景,進展到模擬在其中行動的心智。
English
World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, and considers socially permissible), so a model that tracks the physical scene but not what each agent knows and believes about it predicts the wrong action for the right-looking scene. We formulate Mental World Modeling (MWM), a generic theoretical framework that makes mental variables core components of a world model rather than posthoc rationales: MWM aintains a coupled physical-mental world state, renders a target-specific partial observation, and simulates how candidate actions jointly update both components. We instantiate the framework in MENTIS, a training-free and fully inspectable baseline that decomposes the process into state parsing, target-observation generation, action decomposition, coupled physical and mental transition, and branch-level value evaluation. On a manually constructed, quality-controlled dataset of situated decision scenarios spanning text, image, and sounding-video stories, experiments with 8 modern LLM-based world models demonstrate that explicitly modeling the mental state is essential for predicting human decisions. Deeper analyses further expose the bottlenecks of current mental world modeling. We expect MWM as a next stage of world modeling, from simulating physical scenes to simulating the minds that act in them.