ChatPaper.aiChatPaper

정신적 세계 모델링

Mental World Modeling

July 29, 2026
저자: Hao Fei, Yiran Zhao
cs.AI

초록

세계 모델은 계획과 행동을 위한 예측 기반을 제공하지만, 기존의 공식화는 그것이 무엇이고 어디에 있으며 어떻게 진화할 것인가라는 물리적 질문에만 답한다. 그러나 인간 행동은 은닉된 정신 상태(사람이 무엇을 믿고, 원하고, 의도하고, 느끼며, 사회적으로 허용된다고 간주하는지)에 의해 추동되므로, 물리적 장면을 추적하되 각 에이전트가 그 장면에 대해 알고 믿는 바를 추적하지 않는 모델은 올바르게 보이는 장면에 대해 잘못된 행동을 예측한다. 우리는 정신 세계 모델링(MWM)을 정식화한다. 이는 정신 변수를 사후적 근거가 아닌 세계 모델의 핵심 구성요소로 만드는 일반적 이론적 프레임워크이다. MWM은 결합된 물리-정신 세계 상태를 유지하고, 대상별 부분 관찰을 렌더링하며, 후보 행동이 두 구성요소를 어떻게 함께 갱신하는지를 시뮬레이션한다. 우리는 이 프레임워크를 훈련이 불필요하고 완전히 검사 가능한 베이스라인인 MENTIS에 구현한다. MENTIS는 이 과정을 상태 파싱, 대상 관찰 생성, 행동 분해, 결합된 물리-정신 전이, 분기 수준 가치 평가로 분해한다. 텍스트, 이미지, 음향 비디오 스토리를 아우르는 상황적 의사결정 시나리오로 구성된 수작업 구축, 품질 관리 데이터셋에서 8개의 현대 LLM 기반 세계 모델에 대한 실험은 정신 상태를 명시적으로 모델링하는 것이 인간 의사결정 예측에 필수적임을 입증한다. 더 깊은 분석은 현재 정신 세계 모델링의 병목 지점을 추가로 드러낸다. 우리는 MWM이 세계 모델링의 다음 단계, 즉 물리적 장면을 시뮬레이션하는 것에서 그 장면 안에서 행동하는 마음을 시뮬레이션하는 것으로의 전환을 기대한다.
English
World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, and considers socially permissible), so a model that tracks the physical scene but not what each agent knows and believes about it predicts the wrong action for the right-looking scene. We formulate Mental World Modeling (MWM), a generic theoretical framework that makes mental variables core components of a world model rather than posthoc rationales: MWM aintains a coupled physical-mental world state, renders a target-specific partial observation, and simulates how candidate actions jointly update both components. We instantiate the framework in MENTIS, a training-free and fully inspectable baseline that decomposes the process into state parsing, target-observation generation, action decomposition, coupled physical and mental transition, and branch-level value evaluation. On a manually constructed, quality-controlled dataset of situated decision scenarios spanning text, image, and sounding-video stories, experiments with 8 modern LLM-based world models demonstrate that explicitly modeling the mental state is essential for predicting human decisions. Deeper analyses further expose the bottlenecks of current mental world modeling. We expect MWM as a next stage of world modeling, from simulating physical scenes to simulating the minds that act in them.