超越資料擴展:以表徵為中心的視覺-語言-動作模型持續預訓練

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

August 27, 2026
作者: Senqiao Yang, Chengyao Wang, Yuxin Chen, Zixuan Wang, Longxiang Tang, Haokun Gui, Jinhui Ye, Changsheng Lu, Xiaoyang Wu, Mingkang Zhu, Pengguang Chen, Shu Liu, Zhuotao Tian, Hengshuang Zhao, Bei Yu, Jiaya Jia
cs.AI

摘要

擴展機器人數據對於建構通用視覺-語言-動作(VLA)模型至關重要,然而機器人的軌跡數據比網路規模的圖像-文字數據更難擴展,因為具身數據收集成本高昂且對物理世界的覆蓋稀疏。這使得表徵品質成為核心瓶頸:在固定機器人數據預算下,持續預訓練必須將有限的軌跡轉化為可遷移的視覺-動作知識,而非僅僅擬合動作。我們提出VLAct,一個面向VLA的視覺語言模型(VLM)骨幹網路,在任務特定微調之前,先於廣泛、異構、多本體的機器人數據上進行訓練。VLAct透過VLM先驗保留、多頭連續動作聯合監督以及部分統一的跨本體動作佈局,保留了廣泛的VLM先驗並鼓勵跨本體的共享動作語義,同時允許在微調階段使用任務特定的動作頭。在模擬、真實世界以及未見本體遷移等場景中,VLAct在固定微調協議下持續提升下游效能。在LIBERO-Plus和RoboTwin 2.0上,VLAct超越了包括ABot-M0和LingBot-VLA在內的工業級VLA系統,成功率分別達到82.6%和92.5%。在RoboDojo上,VLAct的成功率在所有策略中排名第六,並在兩項指標上均超越了所有明確指定的世界動作模型(WAM)參賽作品。最值得注意的是,在未見人形本體RoboCasa-GR1上,僅使用20%下游軌跡數據的VLAct便超越了使用完整數據的GR00T-N1.6基線。這些成果完全基於開源數據,且僅需16顆GPU的訓練配置即可達成,顯示以表徵為中心的持續預訓練能在適度的計算預算下帶來極具競爭力的效能,且是VLA進展中獨立於數據擴展之外的重要軸線。
English
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.
PDF842September 1, 2026