小米機器人-1:以超過10萬小時的真實世界軌跡擴展視覺-語言-動作模型
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
July 16, 2026
作者: Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangjun Ye, Wen Ye, Han Zhao, Quanyun Zhou
cs.AI
摘要
我們推出小米機器人1號,這是一個基礎的視覺-語言-動作模型,具備兩項核心能力:(1) 能遵循多樣化的語言指令,在未見過的環境中開箱即用地執行廣泛的行動操作任務;(2) 能以最少微調資料高效適應新穎的下游任務。我們提出兩階段訓練方案,包含預訓練與後訓練。在預訓練階段,我們透過在超過10萬小時的真實世界操作軌跡(以UMI設備收集)上進行訓練,賦予模型廣泛且通用的動作生成能力。關鍵在於,我們開發了一套可擴展的自動標註流程,能以描述場景狀態變化的自然語言註解軌跡片段,為動作學習提供豐富且精準的條件。在後訓練階段,我們旨在將這些能力與機器人本體以及人類自然用以提示機器人的命令式指令進行對齊。大量實驗展現了強大的擴展行為。小米機器人1號在預訓練期間,隨著資料規模與模型大小的增加而持續進步。此擴展行為直接轉移至後訓練階段,更強大的預訓練模型能在未見過的環境中帶來更佳的開箱即用真實機器人表現。此外,小米機器人1號可作為一個強大的機器人基礎策略,在複雜、靈巧的任務上能高效微調,展現高資料效率。在多個模擬基準測試中,小米機器人1號超越了當前最佳方法。值得注意的是,它在RoboCasa365上以57.6%的成功率締造新基準,超越了先前的46.6%最佳成績。此外,它在RoboDojo上取得平均20.07分,大幅優於先前的最佳成績(13.07)。程式碼與模型檢查點將會釋出。專案頁面:https://robotics.xiaomi.com/xiaomi-robotics-1.html
English
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.6% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html