ABot-N1: 汎用視覚言語ナビゲーション基盤モデルを目指して

ABot-N1: Toward a General Visual Language Navigation Foundation Model

July 11, 2026
著者: Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Zedong Chu, Xiaolong Wu, Mu Xu
cs.AI

要旨

視覚言語ナビゲーション基盤モデルは、接地された空間的決定のための深い推論と、多様な具現化タスクに対する広範な汎用性を統合することを目的としています。現在のアプローチは、典型的には観測を直接行動にマッピングするモノリシックなポリシーを通じてこの統合を達成していますが、しばしば座標ドリフトやロングテールセマンティクスの不十分な処理に悩まされています。さらに、これらのブラックボックス的なマッピングは解釈可能性を欠き、汎用性、堅牢性、透明性の同時達成を妨げています。我々は、汎用視覚言語ナビゲーション基盤モデルへの一歩であるABot-N1を提案します。これは、二重視覚言語信号によって導かれる遅速アーキテクチャを介して認知と制御を分離することにより、これらの課題に対処します。より具体的には、低速な視覚言語推論器が明示的なチェーン・オブ・ソート推論を実行し、同時にピクセル目標を生成します。このコンパクトな画像空間アンカーポイントの集合は、ポイントゴール、オブジェクトゴール、POIゴール、指示追従、人物追従などの多様なタスクのための普遍的なインターフェースとして機能します。その後、高速な行動専門家がテキストキューとピクセルガイダンスの両方を活用し、ネイティブ制御周波数で連続的なウェイポイントを生成します。明示的な言語的トレースと対になったピクセル接地アンカーを通じて高レベルの意図と低レベルの制御を橋渡しすることにより、我々のアプローチはシミュレーションと実世界のベンチマークにわたって堅牢で汎用的かつ解釈可能なナビゲーションを保証します。ABot-N1は新たな最先端記録を打ち立て、特に都市規模のナビゲーションにおいて大きな改善をもたらします。POI到着率を35.0%向上(77.3%に)、複雑な屋内および屋外シーンで95.4%/92.9%の成功率(SR)を達成しました。また、オブジェクト到達、人物追従、指示追従タスクにおいて優れた堅牢性を維持します。新しいPoint-Goal/POI-Goalベンチマークがオープンソースとして公開され、都市規模のナビゲーション分野を前進させます。
English
Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet they often suffer from coordinate drift and poor handling of long-tail semantics. Furthermore, these black-box mappings lack interpretability, hindering the simultaneous achievement of generality, robustness, and transparency. We present ABot-N1, a step toward a general Visual Language Navigation foundation model, that addresses these challenges by decoupling cognition from control via a slow-fast architecture guided by dual visual-language signals. More specifically, a slow vision-language reasoner performs explicit Chain-of-Thought reasoning while producing a pixel goal. This compact set of image-space anchor points serves as a universal interface for diverse tasks, including point-goal, object-goal, poi-goal, instruction-following, and person-following. Subsequently, a fast action expert leverages both the textual cues and the pixel guidance to generate continuous waypoints at the native control frequency. By bridging high-level intents and low-level control through pixel-grounded anchors paired with explicit linguistic traces, our approach ensures robust, generalizable, and interpretable navigation across simulation and real-world benchmarks. ABot-N1 establishes new state-of-the-art records, delivering massive gains specifically in urban-scale navigation: boosting POI arrival by 35.0% (to 77.3%) and achieving 95.4%/92.9% SR in complex indoor and outdoor scenes. It also maintains superior robustness across object-reaching, person-following, and instruction-following tasks. New Point-Goal/POI-Goal benchmarks are released as open source to advance the field of urban-scale navigation.
PDF811July 15, 2026