ABot-N1: 일반적인 시각 언어 내비게이션 파운데이션 모델을 향하여

ABot-N1: Toward a General Visual Language Navigation Foundation Model

July 11, 2026
저자: Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Zedong Chu, Xiaolong Wu, Mu Xu
cs.AI

초록

비주얼 언어 내비게이션 기반 모델은 다양한 내장형 작업에 대한 폭넓은 범용성과 함께 근거 기반 공간 결정을 위한 심층 추론을 통합하는 것을 목표로 한다. 현재 접근법은 일반적으로 관찰을 직접 행동에 매핑하는 모놀리식 정책을 통해 이러한 통합을 달성하지만, 종종 좌표 표류와 긴 꼬리(long-tail) 의미론 처리의 부족으로 어려움을 겪는다. 더욱이 이러한 블랙박스 매핑은 해석 가능성이 부족하여 일반성, 강건성 및 투명성을 동시에 달성하는 것을 방해한다. 우리는 이러한 문제를 이중 시각-언어 신호에 의해 안내되는 느림-빠름(slow-fast) 아키텍처를 통해 인지와 제어를 분리함으로써 해결하는, 일반적인 비주얼 언어 내비게이션 기반 모델을 향한 한 걸음인 ABot-N1을 제시한다. 보다 구체적으로, 느린 시각-언어 추론기는 픽셀 목표를 생성하면서 명시적 사고 연쇄(Chain-of-Thought) 추론을 수행한다. 이 이미지 공간 앵커 포인트의 간결한 집합은 점-목표, 객체-목표, 관심 지점-목표, 지시 따르기 및 사람 따라가기를 포함한 다양한 작업을 위한 보편적 인터페이스 역할을 한다. 이후 빠른 행동 전문가는 텍스트 단서와 픽셀 안내를 모두 활용하여 고유 제어 주파수에서 연속적인 경유지를 생성한다. 픽셀 기반 앵커를 명시적 언어 추적과 결합하여 높은 수준의 의도와 낮은 수준의 제어를 연결함으로써, 우리의 접근 방식은 시뮬레이션 및 실제 환경 벤치마크 전반에서 강건하고 일반화 가능하며 해석 가능한 내비게이션을 보장한다. ABot-N1은 새로운 최첨단 기록을 수립하며, 특히 도시 규모 내비게이션에서 큰 성능 향상을 제공한다: 관심 지점 도착률을 35.0% 향상시켜 77.3%에 도달하고, 복잡한 실내 및 실외 장면에서 각각 95.4%와 92.9%의 성공률을 달성한다. 또한 객체 도달, 사람 따라가기 및 지시 따르기 작업 전반에 걸쳐 우수한 강건성을 유지한다. 도시 규모 내비게이션 분야의 발전을 위해 새로운 점-목표 및 관심 지점-목표 벤치마크를 오픈소스로 공개한다.
English
Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet they often suffer from coordinate drift and poor handling of long-tail semantics. Furthermore, these black-box mappings lack interpretability, hindering the simultaneous achievement of generality, robustness, and transparency. We present ABot-N1, a step toward a general Visual Language Navigation foundation model, that addresses these challenges by decoupling cognition from control via a slow-fast architecture guided by dual visual-language signals. More specifically, a slow vision-language reasoner performs explicit Chain-of-Thought reasoning while producing a pixel goal. This compact set of image-space anchor points serves as a universal interface for diverse tasks, including point-goal, object-goal, poi-goal, instruction-following, and person-following. Subsequently, a fast action expert leverages both the textual cues and the pixel guidance to generate continuous waypoints at the native control frequency. By bridging high-level intents and low-level control through pixel-grounded anchors paired with explicit linguistic traces, our approach ensures robust, generalizable, and interpretable navigation across simulation and real-world benchmarks. ABot-N1 establishes new state-of-the-art records, delivering massive gains specifically in urban-scale navigation: boosting POI arrival by 35.0% (to 77.3%) and achieving 95.4%/92.9% SR in complex indoor and outdoor scenes. It also maintains superior robustness across object-reaching, person-following, and instruction-following tasks. New Point-Goal/POI-Goal benchmarks are released as open source to advance the field of urban-scale navigation.
PDF811July 15, 2026