InternVLA-A1.5:统一理解、隱式預見與行動以實現組合泛化
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
July 6, 2026
作者: Haoxiang Ma, Junhao Cai, Xiaoxu Xu, Hao Li, Yuyin Yang, Yang Tian, Jiafei Cao, Hongrui Zhu, Zherui Qiu, Zhaxizhuoma, Yuqiang Yang, Jiaqi Peng, Xueyuan Wei, Yangkun Zhu, Jiahao Jiang, Xing Gao, Hanqing Wang, Feng Yuan, Kailin Li, Xueyue Zhu, Tai Wang, Yan Ding, Jiangmiao Pang, Jia Zeng, Jingjing Zhang, Bowen Zhou, Yao Mu, Chunhua Shen, Weinan Zhang
cs.AI
摘要
机器人操控的统一模型旨在赋予单一策略以预训练视觉语言模型的语义先验,以及通过未来预测学习到的物理动力学。然而,现有设计在实践中往往会侵蚀预训练骨干网络的语义能力,导致异构目标之间的干扰,并在像素空间从头学习未来预测,从而未能充分利用预训练视频生成器的动力学先验。我们提出InternVLA-A1.5,该方法基于原生视觉语言模型骨干构建策略,使其持续在视觉问答和子任务预测上进行训练,并附加一个轻量化的统一专家模块用于连续动作生成。未来预测被重构为潜在查询问题:通过一组可学习的“前瞻令牌”,在冻结的预训练视频生成模型监督下,将任务相关的未来信息浓缩为紧凑的潜在编码,从而使策略继承世界模型的动力学先验,而无需学习像素级生成。推理阶段丢弃视频分支,保持实时控制能力。经过120万机器人交互片段和300万多模态样本的预训练,InternVLA-A1.5在全部六个模拟基准测试中取得最佳综合表现。在真实世界中,保留的语义能力在未见过的指令组合上展现最强的组合泛化性,而两项设计共同支撑了长时域任务的执行。
English
Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In practice, existing designs tend to erode the semantics of the pretrained backbone, suffer interference among heterogeneous objectives, and learn future prediction from scratch in pixel space, leaving the dynamics priors of pretrained video generators unexploited. We present InternVLA-A1.5, which builds the policy on a native VLM backbone that keeps training on VQA and subtask prediction, and attaches a lightweight unified expert for continuous action generation. Future prediction is recast as a latent-querying problem, where a small set of learnable foresight tokens condenses the task-relevant future into a compact latent code under the supervision of a frozen pretrained video generation model, so the policy inherits world-model dynamics priors without ever learning pixel-level generation. The video branch is discarded at inference, keeping real-time control. Pretrained on 1.2M robot episodes and 3M multimodal samples, InternVLA-A1.5 achieves the best overall results on all six simulation benchmarks. In the real world, the preserved semantics deliver the strongest compositional generalization on held-out instruction bindings, and the two designs together sustain long-horizon execution.