ChatPaper.aiChatPaper

PhiZero: 물리적 언어를 중심으로 구축된 세계 모델

PhiZero: A World Model Built Around Physical Language

July 30, 2026
저자: Shuyao Shang, Yuqi Wang, Ruopeng Gao, Xu Chen, Tieniu Tan, Lue Fan, Zhaoxiang Zhang
cs.AI

초록

우리는 물리 언어(physical language), 즉 세계 상태 전이의 간결한 이산 표현을 기반으로 구축된 물리 세계 모델 PhiZero를 소개한다. 기존의 물리 세계 모델들은 일반적으로 미래 비디오를 픽셀 공간에서 직접 예측하여 이면의 세계 역학을 고차원 시각 예측기 내에 암시적으로 남겨 둔다. 인간이 시각 경험에서 예측 구조를 추상화하고 이를 자연어로 체계화하여 명시적 추론을 수행하는 능력에 착안하여, 우리는 실제 환경(in-the-wild) 비디오로부터 자기 지도 학습을 통해 물리 언어를 학습하고 이를 사용하여 물리 세계가 어떻게 진화하는지 명시적으로 추론한다. 이에 따라 PhiZero는 추론 후 렌더링(reason-then-render) 패러다임을 채택한다. 즉, 먼저 미래 세계의 진화를 물리 언어 시퀀스로 추론한 다음, 추론된 전이를 비디오로 렌더링한다. 생성 및 이해 벤치마크에 걸친 광범위한 실험은 PhiZero가 물리적으로 일관된 세계 진화를 모델링하는 능력을 검증한다. 나아가 사실적이고 상호작용적인 세계 모델링, 세밀한 행동 조건부 시뮬레이션, 제로샷 모션 전이에 대한 잠재력을 보여준다.
English
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans' ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.