InternVLA-A1.5: 理解、潜在先見、行動の統合による構成的一般化
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
July 6, 2026
著者: Haoxiang Ma, Junhao Cai, Xiaoxu Xu, Hao Li, Yuyin Yang, Yang Tian, Jiafei Cao, Hongrui Zhu, Zherui Qiu, Zhaxizhuoma, Yuqiang Yang, Jiaqi Peng, Xueyuan Wei, Yangkun Zhu, Jiahao Jiang, Xing Gao, Hanqing Wang, Feng Yuan, Kailin Li, Xueyue Zhu, Tai Wang, Yan Ding, Jiangmiao Pang, Jia Zeng, Jingjing Zhang, Bowen Zhou, Yao Mu, Chunhua Shen, Weinan Zhang
cs.AI
要旨
ロボット操作のための統一モデルは、事前学習済みVLMが持つ意味論的な事前知識と、未来予測を通じて学習される物理ダイナミクスの両方を単一のポリシーに組み込むことを目的としている。実際には、既存の設計は事前学習済みバックボーンの意味論を損ない、異種目的間での干渉を受け、ピクセル空間で未来予測をゼロから学習するため、事前学習済み動画生成モデルのダイナミクス事前知識を活用できていない。本稿ではInternVLA-A1.5を提案する。これは、VQAとサブタスク予測を継続的に学習するネイティブVLMバックボーン上にポリシーを構築し、連続行動生成のための軽量統一エキスパートを付加する。未来予測は潜在クエリ問題として再定義され、少数の学習可能な先見トークンが、凍結済みの事前学習済み動画生成モデルの監督下でタスク関連の未来をコンパクトな潜在コードに圧縮する。これにより、ポリシーはピクセルレベルの生成を学習することなく、世界モデルのダイナミクス事前知識を継承する。動画ブランチは推論時に破棄され、リアルタイム制御を維持する。1.2Mのロボットエピソードと3Mのマルチモーダルサンプルで事前学習されたInternVLA-A1.5は、6つのシミュレーションベンチマークすべてで最高の総合結果を達成する。実世界では、保持された意味論により未見の指示束縛に対して最も強力な構成的一般化を実現し、これら2つの設計により長期実行が持続される。
English
Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In practice, existing designs tend to erode the semantics of the pretrained backbone, suffer interference among heterogeneous objectives, and learn future prediction from scratch in pixel space, leaving the dynamics priors of pretrained video generators unexploited. We present InternVLA-A1.5, which builds the policy on a native VLM backbone that keeps training on VQA and subtask prediction, and attaches a lightweight unified expert for continuous action generation. Future prediction is recast as a latent-querying problem, where a small set of learnable foresight tokens condenses the task-relevant future into a compact latent code under the supervision of a frozen pretrained video generation model, so the policy inherits world-model dynamics priors without ever learning pixel-level generation. The video branch is discarded at inference, keeping real-time control. Pretrained on 1.2M robot episodes and 3M multimodal samples, InternVLA-A1.5 achieves the best overall results on all six simulation benchmarks. In the real world, the preserved semantics deliver the strongest compositional generalization on held-out instruction bindings, and the two designs together sustain long-horizon execution.