ChatPaper.aiChatPaper

Opschaling van Mixture-of-Experts videopretraining voor belichaamde intelligentie

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

July 8, 2026
Auteurs: Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang, Chaoran Feng, Zijing Hu, Chong Bao, Zichen Xi, Yuqi Gan, Weisen Wang, Yanhong Zeng, Qin Zhao, Zifan Shi, Wei Wu, Hao Ouyang, Qiuyu Wang, Shangzhan Zhang, Jiahao Shao, Yipengjing Sun, Liangxiao Hu, Lunke Pan, Nan Xue, Kecheng Zheng, Yinghao Xu, Xing Zhu, Yujun Shen, Ka Leong Cheng
cs.AI

Samenvatting

Ondanks de recente belofte in robotbesturing, hebben videogeneratieve modellen te lijden onder een domeinmismatch vanwege hun primaire focus op contentcreatie. Zo geeft hun ontwerp inherent voorrang aan visuele getrouwheid en creativiteit boven computationele efficiëntie en fysiek realisme. In dit werk presenteren we LingBot-Video, een op DiT gebaseerd videovoor trainingsparadigma dat specifiek is afgestemd op belichaamde intelligentie. Vanuit architectuurperspectief hanteren we het Mixture-of-Experts (MoE)-raamwerk in plaats van een dicht raamwerk om een betere afweging te bereiken tussen modelleercapaciteit en inferentie-efficiëntie, en we weten het vanaf nul op te schalen. Vanuit dataperspectief bouwen we een dataprofileringsengine die standaard internetvideo's uitbreidt met uitgebreid robotgericht beeldmateriaal, waaronder manipulatie, navigatie en egocentrische perspectieven, om het basismodel een intrinsiek begrip van acties en werelddynamiek te geven. Vanuit trainingsperspectief ontwikkelen we een multidimensionaal beloningssysteem om de afstemming op fysieke rationaliteit en taakvoltooiing af te dwingen, verdergaand dan standaardcriteria zoals esthetiek, promptvolging en bewegingsconsistentie. Uitgebreide evaluaties bevestigen de prestaties en efficiëntie ervan als videofoundationmodel. We dragen LingBot-Video bij aan de gemeenschap als het eerste grootschalige, open-source MoE-videofoundationmodel, in een baanbrekende poging om digitale creativiteit en fysieke actuatie te overbruggen.
English
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inherently prioritizes visual fidelity and creativity over computational efficiency and physical realism. In this work, we present LingBot-Video, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence. From the architecture perspective, we adopt the Mixture-of-Experts (MoE), instead of dense, framework to achieve a better trade-off between modeling capacity and inference efficiency, and manage to scale it up from scratch. From the data perspective, we construct a data profiling engine that augments standard internet videos with extensive robot-oriented footage, encompassing manipulation, navigation, and egocentric perspectives, to equip the base model with an intrinsic understanding of actions and world dynamics. From the training perspective, we develop a multi-dimensional reward system to enforce the alignment regarding physical rationality and task completion, going beyond standard criteria such as aesthetics, prompt-following, and motion consistency. Comprehensive evaluations validate its performance and efficiency as a video foundation model. We contribute LingBot-Video as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.