ChatPaper.aiChatPaper

PhiZero:物理言語を中心に構築された世界モデル

PhiZero: A World Model Built Around Physical Language

July 30, 2026
著者: Shuyao Shang, Yuqi Wang, Ruopeng Gao, Xu Chen, Tieniu Tan, Lue Fan, Zhaoxiang Zhang
cs.AI

要旨

我々は、物理言語を中心に構築された物理世界モデルであるPhiZeroを紹介する。物理言語とは、世界状態遷移のコンパクトな離散表現である。既存の物理世界モデルは通常、将来のビデオをピクセル空間で直接予測し、基盤となる世界のダイナミクスを高次元の視覚予測器内に暗黙のうちに残している。人間が視覚経験から予測構造を抽象化し、それを明示的推論のために自然言語で整理する能力に動機づけられ、我々は自己教師あり学習を通じて実世界のビデオから物理言語を学習し、それを用いて物理世界の進化を明示的に推論する。これに基づき、PhiZeroは「推論してからレンダリングする」パラダイムを採用する。すなわち、まず将来の世界進化を物理言語シーケンスとして推論し、次に推論された遷移をビデオにレンダリングする。生成および理解ベンチマークにわたる広範な実験により、PhiZeroが物理的に整合的な世界進化をモデル化する能力が検証される。さらに、現実的かつ対話的な世界モデリング、きめ細かな行動条件付きシミュレーション、ゼロショット動作転移におけるその可能性を示す。
English
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans' ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.