3D HAMSTER: 3D軌道誘導による階層型視覚言語行動モデルにおける計画と制御の橋渡し
3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
June 30, 2026
著者: Dongyoon Hwang, Byungkun Lee, Dongjin Kim, Hyojin Jang, Hoiyeong Jin, Jueun Mun, Minho Park, Hojoon Lee, Hyunseung Kim, Jaegul Choo
cs.AI
要旨
階層的Vision-Language-Action(VLA)モデルは、高レベル計画と低レベル制御を分離することで、ロボット操作における汎化性能を向上させる。このパラダイムにおける最近の研究では、Vision-Language Model(VLM)によって予測された2Dエンドエフェクタ軌道を、下流のポリシに対する明示的なガイダンスとして用いている。しかし、最先端の低レベルポリシは点群上の3Dメトリック空間で動作するため、深度情報を欠く2Dガイダンスを与えると、各ウェイポイントにその下にあるシーン表面の深度が強制的に割り当てられ、幾何学的に歪んだ軌道が生成される。我々は3D HAMSTERを提案する。これは、プランナが直接メトリックに信頼できる3D軌道を出力することで、このギャップを埋める階層的フレームワークである。我々はVLMに専用の深度エンコーダと高密度深度再構成の目的関数を追加し、3Dウェイポイントシーケンスを予測し、それを点群ベースの低レベルポリシに直接統合する。3D軌道予測、シミュレーション、実世界操作において、3D HAMSTERは一貫してプロプライエタリなVLMや2Dガイドのベースラインを上回り、外観変化シフトや未知の言語・空間・視覚条件下で最大の改善を示す。プロジェクトページは https://davian-robotics.github.io/3D_HAMSTER/ で公開されている。
English
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm uses 2D end-effector trajectories predicted by a Vision-Language Model (VLM) as explicit guidance for a downstream policy. However, state-of-the-art low-level policies operate in 3D metric space on point clouds, and feeding them 2D guidance that lacks depth forces each waypoint to be assigned the depth of whatever scene surface lies beneath it, producing geometrically distorted trajectories. We propose 3D HAMSTER, a hierarchical framework that closes this gap by having the planner directly output metrically reliable 3D trajectories. We augment a VLM with a dedicated depth encoder and a dense depth reconstruction objective to predict 3D waypoint sequences, which are directly integrated into a pointcloudbased low-level policy. Across 3D trajectory prediction, simulation, and real-world manipulation, 3D HAMSTER consistently outperforms proprietary VLMs and 2D-guided baselines, with the largest gains under appearance-altering shifts and unseen language, spatial, and visual conditions. The project page is available at https://davian-robotics.github.io/3D_HAMSTER/.