ChatPaper.aiChatPaper

BridgeVLA++: 3次元操作のためのデータ効率的かつ汎化可能なメモリ増強型視覚-言語-行動フレームワーク

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

August 5, 2026
著者: Peiyan Li, Yuze Zhu, Yixiang Chen, Qisen Ma, Yuan Xu, Jiabing Yang, He Guan, Yan Huang, Hongtao Wu, Xiao Ma, Tao Kong, Liang Wang, Tieniu Tan
cs.AI

要旨

事前学習済み視覚-言語モデル(VLM)を活用して視覚-言語-行動(VLA)モデルを構築することは、3次元ロボット操作の有望なパラダイムとして浮上している。しかし、既存の3D VLA手法は依然としてデータ依存性が高く、分布シフト下では限られた汎化しか示さず、過去の観測の明示的な記憶を欠いている。これらの限界は、データが乏しく、オープンワールドで、記憶に依存する操作シナリオへの応用を妨げている。我々の以前の研究であるBridgeVLAは、3次元行動学習中に事前学習済みVLMの入出力の整合性を維持することで、データ効率と汎化を向上させた。すなわち、生の点群を多視点画像に投影し、ロボット行動を生成する前に中間ヒートマップを予測する。本研究では、BridgeVLAに持続的な空間的文脈と時間的相互作用履歴をモデル化する統合時空間メモリアーキテクチャを装備することで、BridgeVLA++を開発する。得られたメモリ拡張フレームワークは、BridgeVLAのデータ効率と汎化能力を維持しながら、観測履歴に基づいて推論することができる。広範な実験により、我々のフレームワークが空間的操作タスクで強力な性能を達成し、頑健な汎化を示すことが実証された。さらにBridgeVLA++は、元のBridgeVLAのデータ効率と汎化性を犠牲にすることなく、2つの困難な記憶依存操作ベンチマークで最先端の性能を達成する。加えて、BridgeVLA++は双腕操作設定でも効果的に動作し、追加の実世界ロボットプラットフォームでも検証されており、タスク、環境、ロボットプラットフォームを横断したスケーラビリティを示している。これらの結果は、BridgeVLA++が、データ効率的な学習、頑健な汎化、効果的な記憶を考慮したロボット操作を同時にサポートする統合3次元視覚-言語-行動フレームワークとして確立されることを示している。プロジェクトウェブサイト: https://bridgevla-plus.github.io/。
English
Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution shifts, and lack explicit memory of past observations. These limitations hinder their application to data-scarce, open-world, and memory-dependent manipulation scenarios. Our previous work, BridgeVLA, improves data efficiency and generalization by preserving the input--output alignment of a pre-trained VLM during 3D action learning: raw point clouds are projected into multi-view images, and intermediate heatmaps are predicted before generating robot actions. In this work, we develop BridgeVLA++ by equipping BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history. The resulting memory-augmented framework can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities. Extensive experiments show that our framework achieves strong performance on spatial manipulation tasks while exhibiting robust generalization. BridgeVLA++ further achieves state-of-the-art performance on two challenging memory-dependent manipulation benchmarks without sacrificing the data efficiency and generalization of the original BridgeVLA. In addition, BridgeVLA++ performs effectively in bimanual manipulation settings and is validated on an additional real-world robotic platform, demonstrating its scalability across tasks, environments, and robotic platforms. These results establish BridgeVLA++ as a unified 3D vision-language-action framework that simultaneously supports data-efficient learning, robust generalization, and effective memory-aware robot manipulation. Project website: https://bridgevla-plus.github.io/.