ChatPaper.aiChatPaper

コードとしての世界:物理的推論のための実行可能な世界表現のエージェントによる発見

Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning

August 27, 2026
著者: Hanyang Wang, Yimo Cai, Weiliang Chen, Jiawei Chi, Haowen Sun, Qiyu Dai, Yi-Hsin Hung, Xingzhuo Guo, Jinshan Ren, Runmao Yao, Ziwei Liu, Mingsheng Long, Yueqi Duan, Jun Gao, Jiangran Lyu, Fangfu Liu, Jialong Wu
cs.AI

要旨

物理的理解と推論は、世界のコンパクトで汎化可能な表現を形成することに依存する。現代の視覚言語モデルは多様な物理的事象を認識・説明できる一方で、世界がどのように進化し介入にどのように応答するかを確実に推論するために必要な、基盤となるメカニズム(物体の状態、物理パラメータ、支配的なダイナミクスなど)の明示的な表現を欠いていることが多い。本研究では、実行可能な世界表現を通じて物理世界を表現するパラダイムであるCode-as-Worldを提案する。物理的構成、動的進化、視覚的外観を実行可能なコードとして表現することにより、Code-as-Worldは物理世界のコンパクトで定量的に根拠づけられ、制御可能な抽象化を提供する。そのような表現を自然言語記述や実世界ビデオなどのマルチモーダル観測から構築するために、我々はアブダクション推論に着想を得たエージェント的発見ループを開発する。このループでは、エージェントが実行可能な世界仮説を提案、実行、レンダリング、検証し、反復的に洗練する。具体的な応用として、我々は検証済みの実行可能世界を用いて、定量的物理推論における視覚言語モデルの訓練のためのスケーラブルな物理的監督を提供する。実験の結果、Code-as-World-VLはQuantiPhyにおいて最先端の性能を達成し、主要なプロプライエタリモデルを上回り、実行可能世界表現が物理的知能のスケーラブルな基盤となる可能性を示している。
English
Physical understanding and reasoning depend on forming compact and generalizable representations of the world. While modern vision-language models can recognize and explain diverse physical events, they often lack explicit representations of the underlying mechanisms-such as object states, physical parameters, and governing dynamics-needed for reliably reasoning how the world evolves and responds to interventions. In this work, we introduce Code-as-World, a paradigm that represents physical worlds through executable world representations. By expressing physical composition, dynamic evolution, and visual appearance as executable code, Code-as-World provides a compact, quantitatively grounded, and controllable abstraction of the physical world. To construct such representations from multimodal observations, such as natural-language descriptions or real-world videos, we develop an agentic discovery loop inspired by abductive reasoning, where an agent proposes, executes, renders, verifies, and iteratively refines executable world hypotheses. As a concrete application, we use verified executable worlds to provide scalable physical supervision for training vision-language models on quantitative physical reasoning. Experiments show that Code-as-World-VL achieves state-of-the-art performance on QuantiPhy and surpasses leading proprietary models, highlighting the potential of executable world representations as a scalable foundation for physical intelligence.