Xiaomi-Robotics-1: 10万時間以上の実世界軌跡データを用いた視覚言語行動モデルのスケーリング

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

July 16, 2026
著者: Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangjun Ye, Wen Ye, Han Zhao, Quanyun Zhou
cs.AI

要旨

以下は日本語訳です。 我々は、Xiaomi-Robotics-1を提案する。これは基盤的な視覚・言語・動作モデルであり、(1)多様な言語指示に従い、未見の環境で幅広いモバイル操作タスクを即座に実行でき、(2)最小限の微調整データで新規の下流タスクに効率的に適応可能である。我々は、事前学習と事後学習からなる2段階の訓練手法を提案する。事前学習では、UMIデバイスにより収集された10万時間以上に及ぶ実世界の操作軌跡データを用いて訓練し、広範かつ汎用的な行動生成能力をモデルに付与する。重要な点として、我々はスケーラブルな自動ラベリングパイプラインを開発し、シーンの状態遷移を記述した自然言語で軌跡クリップに注釈を付与することで、行動学習のための豊富かつ正確な条件付けを提供する。事後学習では、これらの能力をロボットの身体性や、人間がロボットを促すために自然に使用する命令文と整合させることを目指す。大規模な実験により、強いスケーリング特性が示された。Xiaomi-Robotics-1は、事前学習においてデータ規模とモデルサイズの増加に伴い一貫して性能が向上する。このスケーリング特性は事後学習にも直接的に転移し、より強力な事前学習モデルが未見環境での即時的な実ロボット性能の向上をもたらす。さらに、Xiaomi-Robotics-1は強力なロボット基盤方針として機能し、複雑で器用なタスクに対しても高いデータ効率で効率的に微調整できる。複数のシミュレーションベンチマークにおいて、Xiaomi-Robotics-1は最先端手法を凌駕する。特筆すべきは、RoboCasa365において成功率57.6%で新たな最先端記録を達成し、従来の最良値46.6%を上回ったことである。さらに、RoboDojoでは平均スコア20.07を獲得し、従来の最先端値13.07を大きく上回った。コードとモデルチェックポイントは公開予定である。プロジェクトページ: https://robotics.xiaomi.com/xiaomi-robotics-1.html
English
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.6% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html
PDF541July 21, 2026