ChatPaper.aiChatPaper

小米机器人-U0:基于世界基础模型的统一具身综合

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

July 13, 2026
作者: Xinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li
cs.AI

摘要

最近的基础图像和视频生成模型具备强大的泛化能力和可控性,但它们直接应用于具身场景时,受限于多视图一致性、几何一致性以及机器人具身约束等要求。现有方法通常使用有限的机器人数据微调基础模型,往往牺牲了在大规模预训练中获取的视觉知识。我们提出小米机器人-U0(Xiaomi-Robotics-U0),这是一个拥有380亿参数的多模态自回归模型,用于统一的具身合成。它将具身生成视为基础图像和视频生成的扩展,并联合优化了文生图、图像编辑、具身场景生成、具身迁移以及具身视频生成任务。这一统一框架在保留预训练世界基础模型泛化能力的同时,使其适应具身场景。小米机器人-U0是首个支持跨多种机器人具身形态的高质量多视图场景生成模型,并引入了结构化的、可控的具身迁移,实现细粒度编辑,同时保持多视图一致性和交互动态。它在单步和序列生成任务上取得了最先进的结果,在具身场景生成与迁移的人工评估中优于GPT-Image-2.0,在具身视频生成的世界竞技场(World Arena)中排名第一,并在具有挑战性的真实世界操控任务中将pi_0.5的分布外成功率从36.9%提升至63.2%。这些结果表明,基础世界模型既可以作为具身世界模型,也可以作为具身智能的可扩展数据引擎。代码和检查点可通过https://robotics.xiaomi.com/xiaomi-robotics-u0.html获取。
English
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. It treats embodied generation as an extension of foundation image and video generation and jointly optimizes text-to-image generation, image editing, embodied scene generation, embodied transfer, and embodied video generation. This unified framework preserves the generalization of the pre-trained world foundation model while adapting it to embodied settings. Xiaomi-Robotics-U0 is the first model to support high-quality multi-view scene generation across multiple robot embodiments and to introduce structured, controllable embodied transfer for fine-grained editing while preserving multi-view consistency and interaction dynamics. It achieves state-of-the-art results on single-step and sequential generation tasks, outperforming GPT-Image-2.0 in human evaluations of embodied scene generation and transfer, ranking first on World Arena for embodied video generation, and improving the out-of-distribution success rate of pi_0.5 from 36.9% to 63.2% on challenging real-world manipulation tasks. These results show that foundation world models can serve both as embodied world models and scalable data engines for embodied intelligence. Code and checkpoints are available at https://robotics.xiaomi.com/xiaomi-robotics-u0.html.