ABot-N1 : Vers un modèle de fondation général de navigation en langage visuel
ABot-N1: Toward a General Visual Language Navigation Foundation Model
July 11, 2026
Auteurs: Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Zedong Chu, Xiaolong Wu, Mu Xu
cs.AI
Résumé
Les modèles fondamentaux de navigation visuelle-langagère visent à unifier le raisonnement profond pour des décisions spatiales ancrées avec une large polyvalence pour diverses tâches incarnées. Les approches actuelles réalisent généralement cette intégration via des politiques monolithiques qui mappent directement les observations en actions, mais elles souffrent souvent de dérive de coordonnées et d'une mauvaise gestion des sémantiques rares. De plus, ces mappings en boîte noire manquent d'interprétabilité, entravant l'obtention simultanée de généralité, robustesse et transparence. Nous présentons ABot-N1, un pas vers un modèle fondamental de navigation visuelle-langagère général, qui relève ces défis en découplant la cognition du contrôle via une architecture lente-rapide guidée par des signaux visuels-langagiers doubles. Plus précisément, un raisonneur visuel-langagier lent effectue un raisonnement explicite en chaîne de pensée tout en produisant un objectif pixel. Cet ensemble compact de points d'ancrage dans l'espace image sert d'interface universelle pour diverses tâches, incluant l'objectif ponctuel, l'objectif d'objet, l'objectif de point d'intérêt, le suivi d'instructions et le suivi de personne. Ensuite, un expert d'action rapide exploite à la fois les indices textuels et le guidage pixel pour générer des points de passage continus à la fréquence de contrôle native. En reliant les intentions de haut niveau et le contrôle de bas niveau via des points d'ancrage ancrés dans le pixel associés à des traces linguistiques explicites, notre approche assure une navigation robuste, généralisable et interprétable à travers des benchmarks en simulation et dans le monde réel. ABot-N1 établit de nouveaux records de pointe, offrant des gains massifs notamment dans la navigation à l'échelle urbaine : une augmentation de 35,0 % de l'arrivée aux points d'intérêt (à 77,3 %) et un taux de succès de 95,4 %/92,9 % dans des scènes intérieures et extérieures complexes. Il maintient également une robustesse supérieure dans les tâches d'atteinte d'objet, de suivi de personne et de suivi d'instructions. De nouveaux benchmarks d'objectif ponctuel et d'objectif de point d'intérêt sont publiés en open source pour faire avancer le domaine de la navigation à l'échelle urbaine.
English
Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet they often suffer from coordinate drift and poor handling of long-tail semantics. Furthermore, these black-box mappings lack interpretability, hindering the simultaneous achievement of generality, robustness, and transparency. We present ABot-N1, a step toward a general Visual Language Navigation foundation model, that addresses these challenges by decoupling cognition from control via a slow-fast architecture guided by dual visual-language signals. More specifically, a slow vision-language reasoner performs explicit Chain-of-Thought reasoning while producing a pixel goal. This compact set of image-space anchor points serves as a universal interface for diverse tasks, including point-goal, object-goal, poi-goal, instruction-following, and person-following. Subsequently, a fast action expert leverages both the textual cues and the pixel guidance to generate continuous waypoints at the native control frequency. By bridging high-level intents and low-level control through pixel-grounded anchors paired with explicit linguistic traces, our approach ensures robust, generalizable, and interpretable navigation across simulation and real-world benchmarks. ABot-N1 establishes new state-of-the-art records, delivering massive gains specifically in urban-scale navigation: boosting POI arrival by 35.0% (to 77.3%) and achieving 95.4%/92.9% SR in complex indoor and outdoor scenes. It also maintains superior robustness across object-reaching, person-following, and instruction-following tasks. New Point-Goal/POI-Goal benchmarks are released as open source to advance the field of urban-scale navigation.