基礎から応用へ:VLAモデルの実践的改良
From Foundation to Application: Improving VLA Models in Practice
July 7, 2026
著者: Wei Wu, Fangjing Wang, Fan Lu, He Sun, Shi Liu, Yunnan Wang, Yibin Yan, Yong Wang, Shuailei Ma, Xinyang Wang, Yibin Liu, Shuai Yang, Tianxiang Zhou, Kejia Zhang, Lei Zhou, Cheng Su, Nan Xue, Bin Tan, Han Zhang, Youchao Zhang, Fei Liao, Xing Zhu, Yujun Shen, Kecheng Zheng
cs.AI
要旨
最近のVLA基盤モデルの進展にもかかわらず、実験室環境と実世界応用との間の乖離が、その実用的な導入を依然として妨げている。このギャップを埋めるため、我々はLingBot-VLA 2.0を提案する。これはLingBot-VLAを三つの機能領域における改良を通じて発展させたものである。(1) タスクと身体性を横断した汎化。前バージョンと比較し、データ処理パイプラインを刷新し、事前学習用に約6万時間のデータをキュレーションした。その内訳は、20種類のロボット構成にわたる5万時間のロボット軌跡と、1万時間の一人称視点の人間ビデオである。(2) デュアルアームハードウェアプラットフォームに加えた動作空間の拡張。具体的には、本システムは頭部、腰部、移動ベース、器用なハンドの自由度を扱うことで、ロボットが実用的なシナリオにおいてより複雑なタスクに取り組むことを可能にする。(3) 時間的推論を改善するための予測的ダイナミクスモデリング。具体的には、将来予測を代理タスクとして定式化し、意味的先行知識を得るためのビデオ表現モデルと、幾何学的手がかりを得るための深度推定モデルによって支援する。ゼネラリスト設定で実施されたGM-100ベンチマークによる評価は、これらの提案された改良の有益な影響を検証している。さらに、全身自由度をカバーする拡大された事前学習データの恩恵により、LingBot-VLA-2.0は二つのロボットプラットフォームにわたって、強力な身体横断的長期間移動操作能力を示している。
English
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bridge this gap, we present LingBot-VLA 2.0, which advances LingBot-VLA through improvements in three functional domains. (1) Generalization across tasks and embodiments. Compared to the previous version, we revamp the data processing pipeline and curate around 60,000 hours of data for pretraining, including 50,000 hours of robot trajectories spanning 20 robot configurations and 10,000 hours of egocentric human videos. (2) Expanded action space in addition to dual-arm hardware platforms. In particular, our system accommodates degrees of freedom for the heads, waists, mobile bases, and dexterous hands, thereby empowering the robots to tackle more complex tasks in practical scenarios. (3) Predictive dynamics modeling for improved temporal reasoning. Specifically, we formulate future prediction as a proxy task, facilitated by a video representation model for semantic priors and a depth estimation model for geometric cues. Evaluations on the GM-100 benchmark, conducted in a generalist setting, validate the beneficial impact of these proposed modifications. Furthermore, benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.