ChatPaper.aiChatPaper

BridgeVLA++: 3D 조작을 위한 데이터 효율적이고 일반화 가능한 메모리 증강 비전-언어-행동 프레임워크

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

August 5, 2026
저자: Peiyan Li, Yuze Zhu, Yixiang Chen, Qisen Ma, Yuan Xu, Jiabing Yang, He Guan, Yan Huang, Hongtao Wu, Xiao Ma, Tao Kong, Liang Wang, Tieniu Tan
cs.AI

초록

사전 훈련된 비전-언어 모델(VLM)을 활용하여 비전-언어-행동(VLA) 모델을 구축하는 것은 3D 로봇 조작을 위한 유망한 패러다임으로 부상하고 있다. 그러나 기존의 3D VLA 방법들은 여전히 많은 데이터를 요구하며, 분포 변화 하에서 제한된 일반화를 보이고, 과거 관측에 대한 명시적 메모리가 부족하다. 이러한 한계는 데이터가 부족하고 개방형이며 메모리에 의존적인 조작 시나리오에 적용하는 것을 방해한다. 이전 연구인 BridgeVLA는 3D 행동 학습 중 사전 훈련된 VLM의 입력-출력 정렬을 유지함으로써 데이터 효율성과 일반화를 개선한다: 원시 포인트 클라우드를 다중 뷰 이미지로 투영하고, 로봇 행동을 생성하기 전에 중간 히트맵을 예측한다. 본 연구에서는 지속적인 공간 맥락과 시간적 상호작용 이력을 모델링하는 통합 시공간 메모리 아키텍처를 BridgeVLA에 장착하여 BridgeVLA++를 개발한다. 그 결과, 메모리 증강 프레임워크는 BridgeVLA의 데이터 효율성과 일반화 능력을 유지하면서 관측 이력에 대해 추론할 수 있다. 광범위한 실험을 통해 우리 프레임워크가 공간 조작 작업에서 강력한 성능을 달성하고 강건한 일반화를 보여줌을 확인하였다. BridgeVLA++는 또한 원래 BridgeVLA의 데이터 효율성과 일반화를 희생하지 않으면서 두 가지 도전적인 메모리 의존적 조작 벤치마크에서 최첨단 성능을 달성한다. 추가적으로 BridgeVLA++는 양손 조작 환경에서 효과적으로 작동하며, 추가 실제 로봇 플랫폼에서 검증되어 작업, 환경, 로봇 플랫폼에 걸친 확장성을 입증한다. 이러한 결과는 BridgeVLA++가 데이터 효율적 학습, 강건한 일반화, 효과적인 메모리 인지 로봇 조작을 동시에 지원하는 통합 3D 비전-언어-행동 프레임워크임을 확립한다. 프로젝트 웹사이트: https://bridgevla-plus.github.io/.
English
Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution shifts, and lack explicit memory of past observations. These limitations hinder their application to data-scarce, open-world, and memory-dependent manipulation scenarios. Our previous work, BridgeVLA, improves data efficiency and generalization by preserving the input--output alignment of a pre-trained VLM during 3D action learning: raw point clouds are projected into multi-view images, and intermediate heatmaps are predicted before generating robot actions. In this work, we develop BridgeVLA++ by equipping BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history. The resulting memory-augmented framework can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities. Extensive experiments show that our framework achieves strong performance on spatial manipulation tasks while exhibiting robust generalization. BridgeVLA++ further achieves state-of-the-art performance on two challenging memory-dependent manipulation benchmarks without sacrificing the data efficiency and generalization of the original BridgeVLA. In addition, BridgeVLA++ performs effectively in bimanual manipulation settings and is validated on an additional real-world robotic platform, demonstrating its scalability across tasks, environments, and robotic platforms. These results establish BridgeVLA++ as a unified 3D vision-language-action framework that simultaneously supports data-efficient learning, robust generalization, and effective memory-aware robot manipulation. Project website: https://bridgevla-plus.github.io/.