ChatPaper.aiChatPaper

InternVLA-A1.5: 이해, 잠재적 예측, 행동의 통합을 통한 조합적 일반화

InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization

July 6, 2026
저자: Haoxiang Ma, Junhao Cai, Xiaoxu Xu, Hao Li, Yuyin Yang, Yang Tian, Jiafei Cao, Hongrui Zhu, Zherui Qiu, Zhaxizhuoma, Yuqiang Yang, Jiaqi Peng, Xueyuan Wei, Yangkun Zhu, Jiahao Jiang, Xing Gao, Hanqing Wang, Feng Yuan, Kailin Li, Xueyue Zhu, Tai Wang, Yan Ding, Jiangmiao Pang, Jia Zeng, Jingjing Zhang, Bowen Zhou, Yao Mu, Chunhua Shen, Weinan Zhang
cs.AI

초록

로봇 조작을 위한 통합 모델은 사전 학습된 VLM의 의미적 사전 정보와 미래 예측을 통해 학습된 물리적 동역학을 하나의 정책에 탑재하는 것을 목표로 한다. 실제로 기존 설계들은 사전 학습된 백본의 의미를 약화시키고, 이질적인 목표들 간 간섭을 초래하며, 픽셀 공간에서 미래 예측을 처음부터 학습함으로써 사전 학습된 비디오 생성기의 동역학 사전 정보를 활용하지 못하는 경향이 있다. 본 연구는 InternVLA-A1.5를 제시하며, 이는 VQA 및 하위 작업 예측을 지속적으로 학습하는 네이티브 VLM 백본 위에 정책을 구축하고, 연속적인 행동 생성을 위한 경량화된 통합 전문가 모듈을 부착한다. 미래 예측은 잠재 질의 문제로 재정의되며, 소규모의 학습 가능한 예견 토큰(foresight tokens)이 동결된 사전 학습된 비디오 생성 모델의 감독 하에 작업 관련 미래를 간결한 잠재 코드로 압축함으로써, 정책은 픽셀 수준의 생성을 전혀 학습하지 않고도 세계 모델 동역학 사전 정보를 계승한다. 비디오 분기는 추론 시 제거되어 실시간 제어를 유지한다. 120만 개의 로봇 에피소드와 300만 개의 멀티모달 샘플로 사전 학습된 InternVLA-A1.5는 6개 시뮬레이션 벤치마크 전반에서 최고의 종합 성능을 달성한다. 실제 환경에서는 보존된 의미 정보가 보류된 명령 결합에 대해 가장 강력한 조합 일반화를 제공하며, 두 설계가 함께 장기 수행을 유지한다.
English
Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In practice, existing designs tend to erode the semantics of the pretrained backbone, suffer interference among heterogeneous objectives, and learn future prediction from scratch in pixel space, leaving the dynamics priors of pretrained video generators unexploited. We present InternVLA-A1.5, which builds the policy on a native VLM backbone that keeps training on VQA and subtask prediction, and attaches a lightweight unified expert for continuous action generation. Future prediction is recast as a latent-querying problem, where a small set of learnable foresight tokens condenses the task-relevant future into a compact latent code under the supervision of a frozen pretrained video generation model, so the policy inherits world-model dynamics priors without ever learning pixel-level generation. The video branch is discarded at inference, keeping real-time control. Pretrained on 1.2M robot episodes and 3M multimodal samples, InternVLA-A1.5 achieves the best overall results on all six simulation benchmarks. In the real world, the preserved semantics deliver the strongest compositional generalization on held-out instruction bindings, and the two designs together sustain long-horizon execution.