シャオミ・ロボティクスU0: 世界基盤モデルによる統一的身体化合成
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
July 13, 2026
著者: Xinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li
cs.AI
要旨
最近の基盤画像・動画生成モデルは強力な汎化性と制御可能性を提供しているが、それらを具現化シナリオに直接適用することは、多視点一貫性、幾何学的整合性、ロボットの身体性制約といった要件によって制限されている。既存の手法は通常、限られたロボットデータで基盤モデルを適応させることが多く、大規模事前学習で獲得した視覚知識を犠牲にしている。本稿では、統合具現化合成のための380億パラメータのマルチモーダル自己回帰モデル「Xiaomi-Robotics-U0」を提案する。本モデルは、具現化生成を基盤画像・動画生成の拡張として捉え、テキストから画像への生成、画像編集、具現化シーン生成、具現化転移、具現化動画生成を共同最適化する。この統合フレームワークは、事前学習された世界基盤モデルの汎化性を保持しつつ、それを具現化設定に適応させる。Xiaomi-Robotics-U0は、複数のロボット身体性にわたる高品質な多視点シーン生成をサポートし、多視点一貫性と相互作用ダイナミクスを保持しながら、細粒度編集のための構造化された制御可能な具現化転移を導入した最初のモデルである。本モデルは、単一ステップおよび逐次生成タスクにおいて最先端の結果を達成し、具現化シーン生成と転移における人間評価でGPT-Image-2.0を上回り、World Arenaの具現化動画生成で1位を獲得し、挑戦的な実世界操作タスクにおいてpi_0.5の分布外成功率を36.9%から63.2%に改善した。これらの結果は、基盤世界モデルが具現化世界モデルとしても、具現化知能のためのスケーラブルなデータエンジンとしても機能しうることを示している。コードとチェックポイントはhttps://robotics.xiaomi.com/xiaomi-robotics-u0.htmlで公開されている。
English
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. It treats embodied generation as an extension of foundation image and video generation and jointly optimizes text-to-image generation, image editing, embodied scene generation, embodied transfer, and embodied video generation. This unified framework preserves the generalization of the pre-trained world foundation model while adapting it to embodied settings. Xiaomi-Robotics-U0 is the first model to support high-quality multi-view scene generation across multiple robot embodiments and to introduce structured, controllable embodied transfer for fine-grained editing while preserving multi-view consistency and interaction dynamics. It achieves state-of-the-art results on single-step and sequential generation tasks, outperforming GPT-Image-2.0 in human evaluations of embodied scene generation and transfer, ranking first on World Arena for embodied video generation, and improving the out-of-distribution success rate of pi_0.5 from 36.9% to 63.2% on challenging real-world manipulation tasks. These results show that foundation world models can serve both as embodied world models and scalable data engines for embodied intelligence. Code and checkpoints are available at https://robotics.xiaomi.com/xiaomi-robotics-u0.html.