Xiaomi-Robotics-1: 10만 시간 이상의 실제 세계 궤적으로 비전-언어-행동 모델 확장

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

July 16, 2026
저자: Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangjun Ye, Wen Ye, Han Zhao, Quanyun Zhou
cs.AI

초록

저희는 다양한 언어 명령을 따라 처음 보는 환경에서 즉시 다양한 이동 조작 작업을 수행할 수 있으며, 최소한의 미세 조정 데이터로 새로운 하위 과제에 효율적으로 적응할 수 있는 기초 시각-언어-행동(VLA) 모델인 Xiaomi-Robotics-1을 소개합니다. 저희는 사전 학습과 후속 학습으로 구성된 2단계 학습 방식을 제안합니다. 사전 학습 단계에서는 UMI 장치를 통해 수집된 10만 시간 이상의 실제 세계 조작 궤적 데이터로 모델을 훈련하여 광범위하고 일반화 가능한 행동 생성 능력을 부여합니다. 핵심적으로, 장면 상태 전이를 설명하는 자연어로 궤적 클립에 주석을 달아주는 확장 가능한 자동 레이블링 파이프라인을 개발하여 행동 학습에 풍부하고 정확한 조건을 제공합니다. 후속 학습 단계에서는 이러한 능력을 로봇의 물리적 형태 및 인간이 로봇을 지시할 때 자연스럽게 사용하는 명령형 명령과 정렬하는 것을 목표로 합니다. 광범위한 실험을 통해 강력한 스케일링 동작을 확인했습니다. Xiaomi-Robotics-1은 사전 학습 중 데이터 규모와 모델 크기가 증가함에 따라 지속적으로 성능이 향상됩니다. 이러한 스케일링 동작은 후속 학습에 직접적으로 전이되어, 더 강력한 사전 학습 모델이 처음 보는 환경에서 더 나은 즉시 사용 가능한 실제 로봇 성능을 제공합니다. 또한 Xiaomi-Robotics-1은 강력한 로봇 기반 정책으로 작동하여 복잡하고 정밀한 작업에서도 높은 데이터 효율성으로 효율적으로 미세 조정될 수 있습니다. 여러 시뮬레이션 벤치마크에서 Xiaomi-Robotics-1은 최첨단 방법들을 능가합니다. 특히 RoboCasa365에서 57.6%의 성공률로 기존 최고 성능인 46.6%를 크게 상회하는 새로운 최첨단 성과를 달성했습니다. 또한 RoboDojo에서 평균 점수 20.07을 기록하며 이전 최첨단 성과(13.07)를 현저히 능가했습니다. 코드와 모델 체크포인트는 공개될 예정입니다. 프로젝트 페이지: https://robotics.xiaomi.com/xiaomi-robotics-1.html
English
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.6% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html
PDF541July 21, 2026