ChatPaper.aiChatPaper

PhiZero:一個建構於物理語言之上的世界模型

PhiZero: A World Model Built Around Physical Language

July 30, 2026
作者: Shuyao Shang, Yuqi Wang, Ruopeng Gao, Xu Chen, Tieniu Tan, Lue Fan, Zhaoxiang Zhang
cs.AI

摘要

我們提出 PhiZero,一個以物理語言為核心建構的物理世界模型;物理語言是對世界狀態轉換的緊湊離散表徵。現有的物理世界模型通常直接在像素空間中預測未來影片,將底層的世界動態隱含在高維度視覺預測器之中。受到人類能從視覺經驗中抽象出預測結構,並將其組織為自然語言以進行明確推理的能力所啟發,我們透過自我監督從真實世界影片中學習物理語言,並利用它來明確推理物理世界如何演化。因此,PhiZero 採用了「先推理、後渲染」的典範:它首先將未來世界演化推斷為物理語言序列,接著將推斷出的狀態轉換渲染成影片。在生成與理解基準上的大量實驗驗證了 PhiZero 建模物理一致世界演化的能力。我們進一步展示了它在真實且互動的世界建模、細粒度動作條件模擬,以及零樣本動作轉移方面的潛力。
English
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans' ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.