ChatPaper.aiChatPaper

具身導航器:透過指向、思考、記憶與對齊實現高效導航

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

August 18, 2026
作者: Hongyan Feng, Sunlai Chen, Xuanyu Liu, Miao Pan, Yangfan Xie, Yuxiang Cui, Zhongxiang Zhou, Rong Xiong, Wenqi Zhang, Jianwei Yin, Yueting Zhuang, Xuhong Zhang
cs.AI

摘要

儘管大型視覺-語言模型(VLMs)已顯著推動具身導航的發展,其直接部署仍具挑戰性,因為現有方法往往迫使大型視覺-語言模型進入與其二維預訓練先驗不一致的非自然動作空間,再加上僵化的推理排程與低效率的記憶管理。為克服這些限制,我們提出 TAMP-Nav,一個用於高效具身導航的統一框架。首先,我們引入像素到三維動作公式化(Point),將導航重新建構為二維視覺提示。具體而言,大型視覺-語言模型僅需選擇二維像素,再將其投影至三維座標以供底層 SLAM 控制器使用。此設計自然地將具身執行與大型視覺-語言模型固有的二維視覺能力對齊。其次,我們提出整合式選擇性推理與錨定軌跡記憶機制(Think and Memorize),該機制動態觸發鏈式思考,並僅在關鍵節點保留高保真記憶,將冗餘軌跡壓縮為輕量級時空指示器,從而保留關鍵歷史資訊並增強時空感知。最後,我們設計了一個透過群體相對策略優化(GRPO)的高效兩層對齊範式(Align)。透過疊加全局結果獎勵與細粒度過程獎勵,此密集監督將智能體的認知規劃與物理環境回饋緊密對齊,賦予模型自適應推理能力。實驗表明,TAMP-Nav 達到了最先進的效能(例如在 R2R-CE 上成功率達 66.2%),且具有高運行效率和樣本效率(僅需 90k 訓練軌跡)。
English
Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) that reformulates navigation into 2D visual prompting. Specifically, the VLM merely selects 2D pixels, which are then projected into 3D coordinates for a low-level SLAM controller. This design naturally aligns embodied execution with the VLM's inherent 2D visual capabilities. Second, we propose an integrated Selective Reasoning and Anchor-Trajectory Memory mechanism (Think and Memorize), which dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception. Finally, we design an efficient Two-Level Alignment Paradigm (Align) via Group Relative Policy Optimization (GRPO). By superimposing global outcome rewards with fine-grained process rewards, this dense supervision tightly aligns the agent's cognitive planning with physical environmental feedback, endowing the model with adaptive reasoning capabilities. Experiments demonstrate that TAMP-Nav achieves state-of-the-art performance (e.g., 66.2% SR on R2R-CE) with high runtime and sample efficiency (requiring only 90k training trajectories).