3D HAMSTER:通过3D轨迹引导桥接分层视觉语言动作模型中的规划与控制
3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
June 30, 2026
作者: Dongyoon Hwang, Byungkun Lee, Dongjin Kim, Hyojin Jang, Hoiyeong Jin, Jueun Mun, Minho Park, Hojoon Lee, Hyunseung Kim, Jaegul Choo
cs.AI
摘要
層級式視覺-語言-動作(VLA)模型將高層規劃與低層控制解耦,以提升機器人操作的泛化能力。該範疇的近期研究使用視覺-語言模型(VLM)預測的2D末端執行器軌跡,作為下游策略的明確指導。然而,當前最先進的低層策略是在3D度量空間中處理點雲,若餵入缺乏深度資訊的2D引導,每個路徑點必須強制賦予其下方場景表面的深度,導致產生幾何失真的軌跡。我們提出3D HAMSTER,一種層級式架構,透過讓規劃器直接輸出度量可靠的3D軌跡來填補此差距。我們為VLM加上專用深度編碼器與密集深度重建目標,用以預測3D路徑點序列,並將其直接整合至基於點雲的低層策略中。在3D軌跡預測、模擬及真實世界操作任務中,3D HAMSTER持續優於專有VLM與2D引導基線,尤其在外觀變動情境與未見過的語言、空間及視覺條件下獲得最大增益。專案頁面位於 https://davian-robotics.github.io/3D_HAMSTER/。
English
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm uses 2D end-effector trajectories predicted by a Vision-Language Model (VLM) as explicit guidance for a downstream policy. However, state-of-the-art low-level policies operate in 3D metric space on point clouds, and feeding them 2D guidance that lacks depth forces each waypoint to be assigned the depth of whatever scene surface lies beneath it, producing geometrically distorted trajectories. We propose 3D HAMSTER, a hierarchical framework that closes this gap by having the planner directly output metrically reliable 3D trajectories. We augment a VLM with a dedicated depth encoder and a dense depth reconstruction objective to predict 3D waypoint sequences, which are directly integrated into a pointcloudbased low-level policy. Across 3D trajectory prediction, simulation, and real-world manipulation, 3D HAMSTER consistently outperforms proprietary VLMs and 2D-guided baselines, with the largest gains under appearance-altering shifts and unseen language, spatial, and visual conditions. The project page is available at https://davian-robotics.github.io/3D_HAMSTER/.