체화된 지능을 위한 Mixture-of-Experts 비디오 사전 학습 스케일링
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
July 8, 2026
저자: Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang, Chaoran Feng, Zijing Hu, Chong Bao, Zichen Xi, Yuqi Gan, Weisen Wang, Yanhong Zeng, Qin Zhao, Zifan Shi, Wei Wu, Hao Ouyang, Qiuyu Wang, Shangzhan Zhang, Jiahao Shao, Yipengjing Sun, Liangxiao Hu, Lunke Pan, Nan Xue, Kecheng Zheng, Yinghao Xu, Xing Zhu, Yujun Shen, Ka Leong Cheng
cs.AI
초록
최근 로봇 제어 분야에서의 발전에도 불구하고, 비디오 생성 모델은 콘텐츠 생성에 주된 초점을 맞추기 때문에 도메인 불일치 문제를 겪는다. 예를 들어, 이러한 모델의 설계는 본질적으로 계산 효율성과 물리적 현실성보다 시각적 충실도와 창의성을 우선시한다. 본 연구에서는 체화된 지능(embodied intelligence)에 특화된 DiT 기반 비디오 사전 학습 패러다임인 LingBot-Video를 제안한다. 아키텍처 측면에서, 우리는 밀집(dense) 프레임워크 대신 전문가 혼합(MoE) 프레임워크를 채택하여 모델링 용량과 추론 효율성 간의 더 나은 균형을 달성하고, 처음부터 확장하는 데 성공했다. 데이터 측면에서, 우리는 표준 인터넷 비디오에 조작, 내비게이션, 자아 중심 관점을 포함한 광범위한 로봇 중심 영상을 추가로 증강하는 데이터 프로파일링 엔진을 구축하여, 기본 모델에 행동과 세계 역학에 대한 내재적 이해를 갖추게 한다. 훈련 측면에서, 우리는 미학, 프롬프트 준수, 움직임 일관성과 같은 표준 기준을 넘어서, 물리적 합리성과 작업 완료에 관한 정렬을 강화하기 위해 다차원 보상 시스템을 개발한다. 포괄적인 평가를 통해 비디오 기초 모델로서의 성능과 효율성을 검증한다. 우리는 LingBot-Video를 커뮤니티에 최초의 대규모 오픈소스 MoE 비디오 기초 모델로 제공하며, 디지털 창의성과 물리적 작동을 연결하는 선구적인 노력의 일환으로 기여한다.
English
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inherently prioritizes visual fidelity and creativity over computational efficiency and physical realism. In this work, we present LingBot-Video, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence. From the architecture perspective, we adopt the Mixture-of-Experts (MoE), instead of dense, framework to achieve a better trade-off between modeling capacity and inference efficiency, and manage to scale it up from scratch. From the data perspective, we construct a data profiling engine that augments standard internet videos with extensive robot-oriented footage, encompassing manipulation, navigation, and egocentric perspectives, to equip the base model with an intrinsic understanding of actions and world dynamics. From the training perspective, we develop a multi-dimensional reward system to enforce the alignment regarding physical rationality and task completion, going beyond standard criteria such as aesthetics, prompt-following, and motion consistency. Comprehensive evaluations validate its performance and efficiency as a video foundation model. We contribute LingBot-Video as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.