BridgeVLA++:一種資料高效、可泛化且具記憶增強的視覺-語言-動作框架,用於三維操作
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
August 5, 2026
作者: Peiyan Li, Yuze Zhu, Yixiang Chen, Qisen Ma, Yuan Xu, Jiabing Yang, He Guan, Yan Huang, Hongtao Wu, Xiao Ma, Tao Kong, Liang Wang, Tieniu Tan
cs.AI
摘要
利用預訓練視覺-語言模型(VLM)構建視覺-語言-動作(VLA)模型,已成為實現3D機器人操作的一種前景廣闊的範式。然而,現有的3D VLA方法仍高度依賴數據,在分佈偏移下泛化能力有限,且缺乏對過去觀察的顯式記憶。這些局限性阻礙了它們在數據稀缺、開放世界及依賴記憶的操作場景中的應用。我們先前的工作BridgeVLA通過在3D動作學習過程中保持預訓練VLM的輸入-輸出對齊,改善了數據效率與泛化能力:原始點雲被投影為多視圖圖像,並在生成機器人動作之前預測中間熱圖。在本工作中,我們通過為BridgeVLA配備統一的時空記憶架構,對持續的空間上下文與時間交互歷史進行建模,進而開發了BridgeVLA++。由此產生的記憶增強框架能夠在保留BridgeVLA數據效率與泛化能力的同時,對觀察歷史進行推理。大量實驗表明,我們的框架在空間操作任務上表現強勁,同時展現出穩健的泛化能力。BridgeVLA++更進一步在兩個具有挑戰性的依賴記憶的操作基準上達到了最先進的性能,且未犧牲原始BridgeVLA的數據效率與泛化性。此外,BridgeVLA++在雙臂操作設定中亦表現有效,並已在額外的真實機器人平台上得到驗證,展示了其在任務、環境及機器人平台上的可擴展性。這些結果確立了BridgeVLA++作為一個統一的3D視覺-語言-動作框架,同時支持數據高效的學習、穩健的泛化以及有效的記憶感知機器人操作。專案網站:https://bridgevla-plus.github.io/。
English
Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution shifts, and lack explicit memory of past observations. These limitations hinder their application to data-scarce, open-world, and memory-dependent manipulation scenarios. Our previous work, BridgeVLA, improves data efficiency and generalization by preserving the input--output alignment of a pre-trained VLM during 3D action learning: raw point clouds are projected into multi-view images, and intermediate heatmaps are predicted before generating robot actions. In this work, we develop BridgeVLA++ by equipping BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history. The resulting memory-augmented framework can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities. Extensive experiments show that our framework achieves strong performance on spatial manipulation tasks while exhibiting robust generalization. BridgeVLA++ further achieves state-of-the-art performance on two challenging memory-dependent manipulation benchmarks without sacrificing the data efficiency and generalization of the original BridgeVLA. In addition, BridgeVLA++ performs effectively in bimanual manipulation settings and is validated on an additional real-world robotic platform, demonstrating its scalability across tasks, environments, and robotic platforms. These results establish BridgeVLA++ as a unified 3D vision-language-action framework that simultaneously supports data-efficient learning, robust generalization, and effective memory-aware robot manipulation. Project website: https://bridgevla-plus.github.io/.