小米机器人-1:利用超过10万小时的真实世界轨迹数据扩展视觉-语言-动作模型

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

July 16, 2026
作者: Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangjun Ye, Wen Ye, Han Zhao, Quanyun Zhou
cs.AI

摘要

我们提出小米机器人-1,这是一个基础的视觉-语言-动作模型,具备以下能力:(1)在未见环境中开箱即用地遵循多样化的语言指令执行广泛的移动操作任务;(2)以极少的微调数据高效适应新颖的下游任务。我们提出了一种包含预训练和后训练的两阶段训练方案。在预训练阶段,我们通过训练超过10万小时的、由UMI设备收集的真实世界操作轨迹,赋予模型广泛且通用的动作生成能力。关键的是,我们开发了一个可扩展的自动标注流程,用描述场景状态变化的自然语言对轨迹片段进行标注,为动作学习提供了丰富且精确的条件。在后训练阶段,我们旨在将这些能力与机器人本体以及人类自然用于提示机器人的指令性语言对齐。大量实验证明了强大的扩展行为。在预训练中,小米机器人-1的性能随着数据规模和模型大小的增加而持续提升。这种扩展行为直接迁移到后训练中:更强的预训练模型在未见环境中能实现更好的开箱即用真实机器人性能。此外,小米机器人-1作为一个强大的机器人基础策略,能够在复杂、灵巧的任务上以高数据效率进行高效微调。在多个模拟基准测试中,小米机器人-1超越了最先进的方法。值得注意的是,它在RoboCasa365上以57.6%的成功率创下了新的最先进记录,超过了此前的最佳成绩46.6%。同时,它在RoboDojo上的平均得分为20.07,显著优于此前的最先进水平(13.07)。代码和模型检查点将公开发布。项目页面:https://robotics.xiaomi.com/xiaomi-robotics-1.html
English
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.6% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html
PDF541July 21, 2026