ChatPaper.aiChatPaper

面向具身智能的混合專家影片預訓練擴展

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

July 8, 2026
作者: Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang, Chaoran Feng, Zijing Hu, Chong Bao, Zichen Xi, Yuqi Gan, Weisen Wang, Yanhong Zeng, Qin Zhao, Zifan Shi, Wei Wu, Hao Ouyang, Qiuyu Wang, Shangzhan Zhang, Jiahao Shao, Yipengjing Sun, Liangxiao Hu, Lunke Pan, Nan Xue, Kecheng Zheng, Yinghao Xu, Xing Zhu, Yujun Shen, Ka Leong Cheng
cs.AI

摘要

儘管近期在機器人控制方面取得了令人振奮的進展,但影片生成模型因主要專注於內容創作,仍面臨領域不匹配的問題。例如,其設計本質上優先考慮視覺保真度與創造力,而非計算效率與物理真實性。在本研究中,我們提出 LingBot-Video——一種專門為具身智慧量身打造的、基於擴散Transformer(DiT)的影片預訓練範式。從架構角度,我們採用混合專家(MoE)框架而非密集框架,以實現模型容量與推論效率之間更佳的平衡,並成功從零開始將該方法擴展規模。從資料角度,我們建構了一套資料特徵分析引擎,將大量以機器人為導向的影像——涵蓋操作、導航與自我中心視角——融入標準網路影片中,使基礎模型內建對動作與世界動態的固有理解。從訓練角度,我們發展了一套多維獎勵系統,以強化在物理合理性與任務完成度方面的對齊,超越了諸如美學、指令遵循與動作一致性等標準準則。全面的評估驗證了其作為影片基礎模型的性能與效率。我們將 LingBot-Video 貢獻給社群,作為首個大規模、開源的 MoE 影片基礎模型,這是連結數位創意與物理執行的一項開創性努力。
English
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inherently prioritizes visual fidelity and creativity over computational efficiency and physical realism. In this work, we present LingBot-Video, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence. From the architecture perspective, we adopt the Mixture-of-Experts (MoE), instead of dense, framework to achieve a better trade-off between modeling capacity and inference efficiency, and manage to scale it up from scratch. From the data perspective, we construct a data profiling engine that augments standard internet videos with extensive robot-oriented footage, encompassing manipulation, navigation, and egocentric perspectives, to equip the base model with an intrinsic understanding of actions and world dynamics. From the training perspective, we develop a multi-dimensional reward system to enforce the alignment regarding physical rationality and task completion, going beyond standard criteria such as aesthetics, prompt-following, and motion consistency. Comprehensive evaluations validate its performance and efficiency as a video foundation model. We contribute LingBot-Video as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.