ChatPaper.aiChatPaper

Embodied-Navigator:指差し・思考・記憶・整合による効率的なナビゲーション

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

August 18, 2026
著者: Hongyan Feng, Sunlai Chen, Xuanyu Liu, Miao Pan, Yangfan Xie, Yuxiang Cui, Zhongxiang Zhou, Rong Xiong, Wenqi Zhang, Jianwei Yin, Yueting Zhuang, Xuhong Zhang
cs.AI

要旨

大規模視覚言語モデル(VLM)は身体化ナビゲーションを大幅に進歩させてきたが、既存手法はしばしばVLMをその2D事前学習の事前分布と整合しない不自然な行動空間に強制し、さらに固定的な推論スケジュールと非効率的なメモリ管理が重なるため、直接的な展開は依然として困難である。これらの制限を克服するために、我々は効率的な身体化ナビゲーションのための統一フレームワークTAMP-Navを提案する。まず、ナビゲーションを2D視覚プロンプティングとして再定式化するPixel-to-3D Action Formulation(Point)を導入する。具体的には、VLMは2D画素を選択するだけでよく、その画素は低レベルSLAMコントローラ用の3D座標に投影される。この設計により、身体化された実行とVLMが本来持つ2D視覚能力とを自然に整合させることができる。次に、統合型の選択的推論とアンカー軌跡メモリ機構(Think and Memorize)を提案する。これはChain-of-Thoughtを動的にトリガーし、重要ノードにおいてのみ高忠実度メモリを保持する。冗長な軌跡は軽量な時空間指標(Space-Time Indicators)に圧縮され、重要な履歴情報を保持して時空間知覚を向上させる。最後に、Group Relative Policy Optimization(GRPO)による効率的な二段階整合パラダイム(Align)を設計する。大域的な結果報酬ときめ細かい過程報酬を重ね合わせることで、この高密度な監視はエージェントの認知計画と物理的な環境フィードバックとを密接に整合させ、モデルに適応的推論能力を付与する。実験により、TAMP-Navは最先端の性能(例えばR2R-CEで成功率66.2%)を達成し、高い実行時効率とサンプル効率(わずか90kの訓練軌跡のみを必要とする)を示す。
English
Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) that reformulates navigation into 2D visual prompting. Specifically, the VLM merely selects 2D pixels, which are then projected into 3D coordinates for a low-level SLAM controller. This design naturally aligns embodied execution with the VLM's inherent 2D visual capabilities. Second, we propose an integrated Selective Reasoning and Anchor-Trajectory Memory mechanism (Think and Memorize), which dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception. Finally, we design an efficient Two-Level Alignment Paradigm (Align) via Group Relative Policy Optimization (GRPO). By superimposing global outcome rewards with fine-grained process rewards, this dense supervision tightly aligns the agent's cognitive planning with physical environmental feedback, endowing the model with adaptive reasoning capabilities. Experiments demonstrate that TAMP-Nav achieves state-of-the-art performance (e.g., 66.2% SR on R2R-CE) with high runtime and sample efficiency (requiring only 90k training trajectories).