ChatPaper.aiChatPaper

DreamX-Phi 1.0:動作條件化的機器人操控影片世界模型

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

August 13, 2026
作者: DreamX Team, Rui Chen, Xiangxiang Chu, Geng Li, Jifan Li, Qingfeng Shi, Datao Tang, Jing Tang, Jun Wang, Pengfei Zhang
cs.AI

摘要

我們提出DreamX-Phi 1.0,一個用於機器人操作的動作條件化影片世界模型。給定一張觀察幀、一條語言指令,以及由末端執行器姿態與夾爪狀態所組成的指定動作序列,該模型可預測隨之而來的未來觀察結果。然而,僅有真實感並不足以保证忠實度:一個令人信服的生成序列仍可能移動錯誤的機械臂或遺失被操作的物體。為確保預測尊重每個機械臂的指令路徑,我們透過PRoPE風格的幾何編碼,將每個機械臂的SE(3)變換注入注意力機制中,以保留機械臂身份與剛體運動結構。僅靠動作控制無法完全約束場景幾何或小型操作物體的演變。因此,我們加入一個輕量級深度分支來處理場景級幾何,並使用SAM3遮罩搭配凍結的V-JEPA教師模型,在抓取過程中維持物體一致性。我們進一步透過分佈匹配蒸餾,將多步生成器蒸餾為少步學生模型,以實現高效部署。截至撰寫本文時,DreamX-Phi 1.0在WorldArena 2.0挑戰賽的Track 1中獲得第一名,並在Track 2中獲得第二名。我們的模型與程式碼將公開提供。
English
We present DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm SE(3) transformations into attention via PRoPE-style geometric encoding, preserving arm identity and rigid-motion structure. Action control alone does not fully constrain scene geometry or the evolution of small manipulated objects. We therefore add a lightweight depth branch for scene-level geometry and use SAM3 masks with a frozen V-JEPA teacher to maintain object consistency throughout grasping. We further distill the multi-step generator into a few-step student via distribution-matching distillation for efficient deployment. At the time of writing, achieves first place on Track~1 and second place on Track~2 of the WorldArena~2.0 Challenge. Our model and code will be publicly available.