데이터 스케일링을 넘어서: 비전-언어-행동 모델을 위한 표현 중심의 지속적 사전학습

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

August 27, 2026
저자: Senqiao Yang, Chengyao Wang, Yuxin Chen, Zixuan Wang, Longxiang Tang, Haokun Gui, Jinhui Ye, Changsheng Lu, Xiaoyang Wu, Mingkang Zhu, Pengguang Chen, Shu Liu, Zhuotao Tian, Hengshuang Zhao, Bei Yu, Jiaya Jia
cs.AI

초록

로봇 데이터 확장은 범용 비전-언어-행동(VLA) 모델을 구축하는 데 필수적이지만, 로봇 궤적 데이터는 체화 기반 수집 비용이 높고 물리적 세계를 희소하게만 커버하므로 웹 규모의 이미지-텍스트 데이터보다 확장하기 어렵다. 이에 따라 표현 품질이 핵심 병목으로 부상한다: 고정된 로봇 데이터 예산 하에서 지속적 사전 학습은 제한된 궤적을 단순히 행동에 맞추는 것을 넘어, 전이 가능한 시각-행동 지식으로 전환해야 한다. 본 논문은 작업별 미세 조정 전에 광범위하고 이질적인 다중 체화 로봇 데이터로 사전 학습되는 VLA 지향 VLM 백본인 VLAct를 제안한다. VLAct는 VLM 사전 지식 보존, 다중 헤드 연속 행동 공동 지도, 부분적으로 통합된 교차 체화 행동 레이아웃을 통해 광범위한 VLM 사전 지식을 유지하고 체화 간 공유 행동 의미론을 장려하며, 미세 조정 단계에서는 작업별 행동 헤드를 허용한다. 시뮬레이션, 실제 환경, 미경험 체화 전이 전반에 걸쳐 VLAct는 고정된 미세 조정 프로토콜 하에서 다운스트림 성능을 일관되게 향상시킨다. LIBERO-Plus와 RoboTwin 2.0에서 VLAct는 ABot-M0 및 LingBot-VLA를 포함한 산업용 VLA 시스템을 능가하며 각각 82.6%와 92.5%의 성공률을 달성한다. RoboDojo에서 VLAct는 성공률 기준 전체 정책 중 6위를 기록하고, 명시적으로 지정된 모든 세계-행동 모델(WAM) 항목을 두 지표 모두에서 능가한다. 가장 주목할 만한 점은, 미경험 휴머노이드 체화인 RoboCasa-GR1에서 VLAct가 다운스트림 궤적의 20%만 사용하고도 전체 데이터를 사용한 GR00T-N1.6 기준선을 능가한다는 것이다. 이러한 결과는 완전히 오픈소스로 공개된 데이터와 16-GPU 학습 환경만으로 달성되었으며, 이는 표현 중심의 지속적 사전 학습이 소규모 컴퓨팅 예산 하에서도 매우 경쟁력 있는 성능을 제공할 수 있고, 데이터 확장을 넘어서는 VLA 발전의 중요한 독립 축임을 시사한다.
English
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.
PDF842September 1, 2026