データスケーリングを超えて:視覚言語行動モデルのための表現中心の継続的事前学習

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

August 27, 2026
著者: Senqiao Yang, Chengyao Wang, Yuxin Chen, Zixuan Wang, Longxiang Tang, Haokun Gui, Jinhui Ye, Changsheng Lu, Xiaoyang Wu, Mingkang Zhu, Pengguang Chen, Shu Liu, Zhuotao Tian, Hengshuang Zhao, Bei Yu, Jiaya Jia
cs.AI

要旨

ロボットデータのスケーリングは、汎用性の高いVision-Language-Action(VLA)モデルを構築する上で極めて重要である。しかし、ロボット軌道は、身体性を伴う収集にコストがかかり、物理世界を疎にしかカバーしないため、ウェブ規模の画像テキストデータよりもスケーリングが困難である。このことが表現品質を中心的なボトルネックにしている。すなわち、固定されたロボットデータ予算の下では、継続的事前学習は、限られた軌道を単に動作に適合させるのではなく、転移可能な視覚行動知識へと変換しなければならない。我々は、タスク固有のファインチューニングの前に、広範かつ異種混在のマルチ身体性ロボットデータで学習されたVLA指向のVLMバックボーンであるVLActを提案する。VLActは、VLM事前知識の保持、マルチヘッド連続動作の共同監視、および部分的に統合された身体性横断的動作レイアウトにより、広範なVLM事前知識を保ちながら身体性を超えた共有動作セマンティクスを促進する。一方で、ファインチューニング時にはタスク固有のアクションヘッドを許容する。シミュレーション、実世界、および未見の身体性への転移の全体にわたって、VLActは固定されたファインチューニングプロトコルの下で下流タスクのパフォーマンスを一貫して向上させる。LIBERO-PlusおよびRoboTwin 2.0において、VLActはABot-M0やLingBot-VLAを含む産業用VLAシステムを凌駕し、82.6%および92.5%の成功率を達成する。RoboDojoでは、VLActは成功率において全ポリシー中6位にランクインし、明示的に指定されたワールドアクションモデル(WAM)の全エントリを両指標で上回る。最も注目すべきは、未見のヒューマノイド身体性であるRoboCasa-GR1において、VLActが下流タスクの軌道のわずか20%のみを用いて、全データを使用したGR00T-N1.6ベースラインを上回ることである。これらの結果は、完全にオープンソースのデータとわずか16-GPUの学習環境のみを用いて得られており、表現中心の継続的事前学習が控えめな計算予算の下で非常に競争力のあるパフォーマンスを発揮できることを示している。これは、データスケーリングを超えたVLA進歩の重要な独立した軸である。
English
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.
PDF842September 1, 2026