Xiaomi-Robotics-U0: 세계 기반 모델을 통한 통합 체화 합성
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
July 13, 2026
저자: Xinghang Li, Jun Guo, Qiwei Li, Long Qian, Hang Lai, Yueze Wang, Hongyu Yan, Jiahang Cao, Xi Chen, Jingen Qu, Jiaxi Song, Nan Sun, Hanye Zhao, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Caoyu Xia, Jack Zhao, Diyun Xiang, Hangjun Ye, Heng Qu, Huaping Liu, Jason Li
cs.AI
초록
최근의 기반 이미지 및 비디오 생성 모델은 강력한 일반화와 제어 가능성을 제공하지만, 다중 뷰 일관성, 기하학적 일관성, 로봇 체화 제약 조건에 대한 요구로 인해 체화된 시나리오에 직접 적용하는 데 제한이 있습니다. 기존 방법들은 일반적으로 제한된 로봇 데이터로 기반 모델을 적응시키며, 대규모 사전 학습 중에 획득한 시각적 지식을 종종 희생합니다. 우리는 통합 체화 합성을 위한 380억 개의 매개변수를 가진 다중 모드 자기회귀 모델인 Xiaomi-Robotics-U0를 제시합니다. 이 모델은 체화 생성을 기반 이미지 및 비디오 생성의 확장으로 간주하고, 텍스트-이미지 생성, 이미지 편집, 체화 장면 생성, 체화 전이, 체화 비디오 생성을 공동으로 최적화합니다. 이 통합 프레임워크는 사전 학습된 세계 기반 모델의 일반화를 보존하면서 체화 환경에 적응시킵니다. Xiaomi-Robotics-U0는 여러 로봇 체화에 걸쳐 고품질 다중 뷰 장면 생성을 지원하고, 다중 뷰 일관성과 상호작용 역학을 보존하면서 세밀한 편집을 위한 구조화되고 제어 가능한 체화 전이를 도입한 최초의 모델입니다. 이 모델은 단일 단계 및 순차 생성 작업에서 최첨단 결과를 달성하여, 체화 장면 생성 및 전이에 대한 인간 평가에서 GPT-Image-2.0을 능가하고, 체화 비디오 생성에 대한 World Arena에서 1위를 차지했으며, 까다로운 실제 조작 작업에서 pi_0.5의 분포 외 성공률을 36.9%에서 63.2%로 향상시켰습니다. 이러한 결과는 기반 세계 모델이 체화된 세계 모델이자 체화 지능을 위한 확장 가능한 데이터 엔진으로 기능할 수 있음을 보여줍니다. 코드와 체크포인트는 https://robotics.xiaomi.com/xiaomi-robotics-u0.html에서 확인할 수 있습니다.
English
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training. We present Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. It treats embodied generation as an extension of foundation image and video generation and jointly optimizes text-to-image generation, image editing, embodied scene generation, embodied transfer, and embodied video generation. This unified framework preserves the generalization of the pre-trained world foundation model while adapting it to embodied settings. Xiaomi-Robotics-U0 is the first model to support high-quality multi-view scene generation across multiple robot embodiments and to introduce structured, controllable embodied transfer for fine-grained editing while preserving multi-view consistency and interaction dynamics. It achieves state-of-the-art results on single-step and sequential generation tasks, outperforming GPT-Image-2.0 in human evaluations of embodied scene generation and transfer, ranking first on World Arena for embodied video generation, and improving the out-of-distribution success rate of pi_0.5 from 36.9% to 63.2% on challenging real-world manipulation tasks. These results show that foundation world models can serve both as embodied world models and scalable data engines for embodied intelligence. Code and checkpoints are available at https://robotics.xiaomi.com/xiaomi-robotics-u0.html.