小米機器人-U0:基於世界基礎模型的統一具身合成
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
July 13, 2026
作者: Xinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li
cs.AI
摘要
近期推出的基礎影像與影片生成模型雖具備強大的泛化能力與可控性,但其直接應用於具身場景時,仍受限於多視角一致性、幾何連貫性及機器人實體約束等要求。現有方法通常以有限的機器人資料微調基礎模型,卻往往犧牲大規模預訓練階段所習得的視覺知識。我們提出「小米機器人-U0」,這是一個擁有380億參數的多模態自回歸模型,專為統一化的具身合成而設計。該模型將具身生成視為基礎影像與影片生成之延伸,並共同優化文字轉影像生成、影像編輯、具身場景生成、具身遷移以及具身影片生成等任務。此統一框架在保留預訓練世界基礎模型泛化能力的同時,將其適應至具身設定。小米機器人-U0是首個能支援多種機器人實體下高品質多視角場景生成的模型,並引入結構化、可控的具身遷移技術,實現細粒度編輯,同時維持多視角一致性與互動動態。在單步生成與序列生成任務中,該模型達到業界最佳成果:於具身場景生成與遷移的人類評估中勝過GPT-Image-2.0,在World Arena具身影片生成排行榜上名列第一,並將pi_0.5在具挑戰性的真實世界操作任務中,其分佈外成功率從36.9%提升至63.2%。這些結果顯示,基礎世界模型既能作為具身世界模型,也能作為具身智慧的可擴展資料引擎。程式碼與檢查點已公開於 https://robotics.xiaomi.com/xiaomi-robotics-u0.html。
English
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. It treats embodied generation as an extension of foundation image and video generation and jointly optimizes text-to-image generation, image editing, embodied scene generation, embodied transfer, and embodied video generation. This unified framework preserves the generalization of the pre-trained world foundation model while adapting it to embodied settings. Xiaomi-Robotics-U0 is the first model to support high-quality multi-view scene generation across multiple robot embodiments and to introduce structured, controllable embodied transfer for fine-grained editing while preserving multi-view consistency and interaction dynamics. It achieves state-of-the-art results on single-step and sequential generation tasks, outperforming GPT-Image-2.0 in human evaluations of embodied scene generation and transfer, ranking first on World Arena for embodied video generation, and improving the out-of-distribution success rate of pi_0.5 from 36.9% to 63.2% on challenging real-world manipulation tasks. These results show that foundation world models can serve both as embodied world models and scalable data engines for embodied intelligence. Code and checkpoints are available at https://robotics.xiaomi.com/xiaomi-robotics-u0.html.