超越数据扩展:面向视觉-语言-动作模型的表征中心持续预训练
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
August 27, 2026
作者: Senqiao Yang, Chengyao Wang, Yuxin Chen, Zixuan Wang, Longxiang Tang, Haokun Gui, Jinhui Ye, Changsheng Lu, Xiaoyang Wu, Mingkang Zhu, Pengguang Chen, Shu Liu, Zhuotao Tian, Hengshuang Zhao, Bei Yu, Jiaya Jia
cs.AI
摘要
扩展机器人数据对于构建通用视觉-语言-动作(VLA)模型至关重要,但机器人轨迹比网络规模的图像-文本数据更难扩展,因为具身数据采集成本高昂且对物理世界的覆盖稀疏。这使得表征质量成为核心瓶颈:在固定的机器人数据预算下,持续预训练必须将有限的轨迹转化为可迁移的视觉-动作知识,而不仅仅是拟合动作序列。
我们提出了VLAct,一个面向VLA的视觉语言模型骨干网络,在广泛、异质、多具身的机器人数据上进行训练,然后进行任务特化微调。VLAct通过VLM先验保持、多头连续动作协同监督和部分统一的跨具身动作布局,保留了广泛的VLM先验并鼓励跨具身的共享动作语义,同时允许在微调阶段使用任务特定的动作头。在仿真、真实世界和未见具身迁移场景中,VLAct在固定微调协议下持续提升下游性能。在LIBERO-Plus和RoboTwin 2.0上,VLAct超越了包括ABot-M0和LingBot-VLA在内的工业级VLA系统,分别达到82.6%和92.5%的成功率。在RoboDojo上,VLAct的成功率在所有策略中排名第六,并在两项指标上均优于所有明确指定的世界动作模型(WAM)参赛方案。
最值得注意的是,在未见人形具身RoboCasa-GR1上,仅使用20%下游轨迹的VLAct超越了使用全量数据的GR00T-N1.6基线。这些结果完全基于开源数据,且仅使用16块GPU的训练配置,表明以表征为中心的持续预训练能够在适度的计算预算下取得极具竞争力的性能,是VLA进展中超越数据扩展的重要独立维度。
English
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.