具現化知能のための混合専門家動画事前学習の大規模化
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
July 8, 2026
著者: Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang, Chaoran Feng, Zijing Hu, Chong Bao, Zichen Xi, Yuqi Gan, Weisen Wang, Yanhong Zeng, Qin Zhao, Zifan Shi, Wei Wu, Hao Ouyang, Qiuyu Wang, Shangzhan Zhang, Jiahao Shao, Yipengjing Sun, Liangxiao Hu, Lunke Pan, Nan Xue, Kecheng Zheng, Yinghao Xu, Xing Zhu, Yujun Shen, Ka Leong Cheng
cs.AI
要旨
ロボット制御における最近の期待にもかかわらず、動画生成モデルは主にコンテンツ生成に焦点を当てているため、ドメインミスマッチの問題を抱えている。例えば、その設計は本質的に、計算効率や物理的リアリズムよりも視覚的忠実性と創造性を優先している。本研究では、具現化知能に特化したDiTベースの動画事前学習パラダイムであるLingBot-Videoを提案する。アーキテクチャの観点からは、密なフレームワークではなくMixture-of-Experts (MoE)を採用することで、モデリング能力と推論効率のより良いトレードオフを実現し、ゼロからスケールアップすることに成功した。データの観点からは、標準的なインターネット動画に、操作、ナビゲーション、自己中心視点を含む広範なロボット向け映像を拡充するデータプロファイリングエンジンを構築し、基本モデルに行動と世界のダイナミクスに対する内在的理解を備えさせる。訓練の観点からは、美的品質、プロンプト追従、動作一貫性といった標準的な基準を超え、物理的合理性とタスク完了に関する整合性を強制する多次元報酬システムを開発する。包括的な評価により、動画基盤モデルとしての性能と効率が検証された。我々は、デジタル創造性と物理的行動を橋渡しする先駆的取り組みとして、初の大規模オープンソースMoE動画基盤モデルであるLingBot-Videoをコミュニティに提供する。
English
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inherently prioritizes visual fidelity and creativity over computational efficiency and physical realism. In this work, we present LingBot-Video, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence. From the architecture perspective, we adopt the Mixture-of-Experts (MoE), instead of dense, framework to achieve a better trade-off between modeling capacity and inference efficiency, and manage to scale it up from scratch. From the data perspective, we construct a data profiling engine that augments standard internet videos with extensive robot-oriented footage, encompassing manipulation, navigation, and egocentric perspectives, to equip the base model with an intrinsic understanding of actions and world dynamics. From the training perspective, we develop a multi-dimensional reward system to enforce the alignment regarding physical rationality and task completion, going beyond standard criteria such as aesthetics, prompt-following, and motion consistency. Comprehensive evaluations validate its performance and efficiency as a video foundation model. We contribute LingBot-Video as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.