ABot-N1:邁向通用視覺語言導航基礎模型
ABot-N1: Toward a General Visual Language Navigation Foundation Model
July 11, 2026
作者: Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Zedong Chu, Xiaolong Wu, Mu Xu
cs.AI
摘要
视觉语言导航基础模型旨在将深度推理与具身空间决策统一起来,同时保持广泛通用性以应对多样化的具身任务。当前方法通常通过将观测直接映射到动作的单一策略来实现这种整合,但常受坐标漂移和长尾语义处理不佳的困扰。此外,这些黑箱映射缺乏可解释性,难以同时实现通用性、鲁棒性和透明性。我们提出ABot-N1,朝向通用视觉语言导航基础模型迈进,通过双视觉-语言信号引导的慢快架构将认知与控制解耦,从而解决上述挑战。具体而言,慢速视觉-语言推理器执行显式链式思维推理,同时生成像素目标。这一紧凑的图像空间锚点集合作为通用接口,支持点目标、物体目标、兴趣点目标、指令跟随和人物跟随等多样任务。随后,快速动作专家利用文本线索和像素引导,以原生控制频率生成连续航点。通过像素级锚点桥接高层意图与底层控制,并辅以显式语言痕迹,我们的方法在仿真和真实世界基准测试中确保了鲁棒、可泛化且可解释的导航。ABot-N1确立了新的最先进纪录,尤其在城市场景导航中实现了显著提升:兴趣点到达率提升35.0%(达到77.3%),并在复杂室内外场景中取得了95.4%/92.9%的成功率。此外,在物体到达、人物跟随和指令跟随任务中保持了卓越的鲁棒性。我们开源了新的点目标/兴趣点目标基准测试,以推动城市场景导航领域的发展。
English
Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet they often suffer from coordinate drift and poor handling of long-tail semantics. Furthermore, these black-box mappings lack interpretability, hindering the simultaneous achievement of generality, robustness, and transparency. We present ABot-N1, a step toward a general Visual Language Navigation foundation model, that addresses these challenges by decoupling cognition from control via a slow-fast architecture guided by dual visual-language signals. More specifically, a slow vision-language reasoner performs explicit Chain-of-Thought reasoning while producing a pixel goal. This compact set of image-space anchor points serves as a universal interface for diverse tasks, including point-goal, object-goal, poi-goal, instruction-following, and person-following. Subsequently, a fast action expert leverages both the textual cues and the pixel guidance to generate continuous waypoints at the native control frequency. By bridging high-level intents and low-level control through pixel-grounded anchors paired with explicit linguistic traces, our approach ensures robust, generalizable, and interpretable navigation across simulation and real-world benchmarks. ABot-N1 establishes new state-of-the-art records, delivering massive gains specifically in urban-scale navigation: boosting POI arrival by 35.0% (to 77.3%) and achieving 95.4%/92.9% SR in complex indoor and outdoor scenes. It also maintains superior robustness across object-reaching, person-following, and instruction-following tasks. New Point-Goal/POI-Goal benchmarks are released as open source to advance the field of urban-scale navigation.