ABot-N1:迈向通用视觉语言导航基础模型
ABot-N1: Toward a General Visual Language Navigation Foundation Model
July 11, 2026
作者: Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Zedong Chu, Xiaolong Wu, Mu Xu
cs.AI
摘要
视觉语言导航基础模型旨在将深层推理能力与空间决策能力相统一,同时兼具广泛适应性以应对多样的具身任务。当前方法通常通过将观测直接映射为动作的单一策略来实现这一整合,但这类方法常面临坐标漂移以及对长尾语义处理不佳的问题。此外,这些黑箱映射缺乏可解释性,阻碍了通用性、鲁棒性与透明性的同步实现。我们提出ABot-N1,向通用视觉语言导航基础模型迈进一步。该方法通过双视觉-语言信号引导的慢-快架构,将认知与控制解耦,从而解决了上述挑战。具体而言,一个慢速视觉-语言推理器执行显式的思维链推理,同时生成一个像素目标。这一紧凑的图像空间锚点集合充当多种任务的通用接口,包括点目标、对象目标、兴趣点目标、指令跟随以及人物跟随。随后,一个快速动作专家利用文本线索和像素指引,以原生控制频率生成连续路径点。通过将高层意图与低层控制经由像素级锚点及显式语言轨迹相连接,我们的方法在仿真与真实世界基准测试中实现了鲁棒、通用且可解释的导航。ABot-N1刷新了多项最先进纪录,尤其是在城市场景导航中表现突出:兴趣点到达率提升35.0%(达到77.3%),在复杂室内外场景中分别达到95.4%和92.9%的成功率。同时,在物体到达、人物跟随和指令跟随任务中保持了卓越的鲁棒性。我们开源了新的点目标/兴趣点目标基准数据集,以推动城市场景导航领域的发展。
English
Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet they often suffer from coordinate drift and poor handling of long-tail semantics. Furthermore, these black-box mappings lack interpretability, hindering the simultaneous achievement of generality, robustness, and transparency. We present ABot-N1, a step toward a general Visual Language Navigation foundation model, that addresses these challenges by decoupling cognition from control via a slow-fast architecture guided by dual visual-language signals. More specifically, a slow vision-language reasoner performs explicit Chain-of-Thought reasoning while producing a pixel goal. This compact set of image-space anchor points serves as a universal interface for diverse tasks, including point-goal, object-goal, poi-goal, instruction-following, and person-following. Subsequently, a fast action expert leverages both the textual cues and the pixel guidance to generate continuous waypoints at the native control frequency. By bridging high-level intents and low-level control through pixel-grounded anchors paired with explicit linguistic traces, our approach ensures robust, generalizable, and interpretable navigation across simulation and real-world benchmarks. ABot-N1 establishes new state-of-the-art records, delivering massive gains specifically in urban-scale navigation: boosting POI arrival by 35.0% (to 77.3%) and achieving 95.4%/92.9% SR in complex indoor and outdoor scenes. It also maintains superior robustness across object-reaching, person-following, and instruction-following tasks. New Point-Goal/POI-Goal benchmarks are released as open source to advance the field of urban-scale navigation.