LightNav-0: 범용 체화 내비게이션을 위한 VLM 공간 지능 이끌어내기
LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation
August 31, 2026
저자: Shaoan Wang, Aocheng Luo, Fei Huang, Jingyi Xu, Xiaoyang Wang, Yueyu Wang, Qianli Ma, Fan Yang, Ran Mei, Jia Wei, Jiangpeng Hu, Xuhao Liu, Hongming Chen, Yuanbin Shao, Yiyang Lin, Ziliang Li, Liang Pan, Xinhang Liu, Yuntao Ma, Tingxiang Fan
cs.AI
초록
구현 내비게이션은 에이전트가 이질적인 목표와 시각적 관찰을 작업, 환경, 로봇 구현체 전반에 걸친 행동으로 변환하는 것을 요구한다. 현대 비전-언어 모델(VLM)은 이미 시각적 접지, 공간 추론, 포인팅을 위한 공간 사전 지식을 인코딩하지만, 이러한 능력은 로봇 제어를 위해 직접적으로 활용되는 경우가 드물다. 기존 내비게이션 시스템은 대신 작업 또는 구현체 특정 구성요소에 의존하여 지각, 추론, 행동을 분절시키고 제한된 일반화만을 제공한다. 본 연구에서는 사전 훈련된 VLM의 공간 지능을 이끌어 내비게이션에 정렬시키는, 작업 특정 예측 헤드 없이 동작하는 경량 범용 구현 내비게이션 모델 LightNav-0를 제시한다. LightNav-0는 통합 토큰 인터페이스를 통해 다양한 내비게이션 작업을 표현한다: 이중 채널 포인팅은 작업, 장면, 구현체에 무관한 공간 의도를 표현하고, 잔차 벡터 양자화 동작 토크나이저는 이 의도를 정밀한 구현체 특정 궤적으로 변환한다. 시간 인식적 시각 이력 압축, ER(구현 추론) 중간 훈련, 지도 미세 조정, 강화 학습과 함께, 이 구성은 단일 모델 내에서 지시 따르기, 개방형 어휘 객체 내비게이션, 시각적 추적을 지원한다. 내비게이션 훈련 코퍼스는 2K+ 장면과 4K+ 시간의 구현 내비게이션 데이터로 구성된다. LightNav-0 초기화에 사용되는 구현 추론 체크포인트인 LightNav-ER은 8개 구현 추론 벤치마크에서 최고의 전체 세트 평균을 달성하며, LightNav-0는 10개 공개 내비게이션 시뮬레이션 설정 모두에서 최첨단 단안 성공률을 달성한다. 실세계 평가는 로봇 구현체, 다양한 장면, 정적 및 동적 대상에 걸친 제로샷 일반화를 추가로 입증한다. 이러한 결과는 경량 VLM이 범용 구현 내비게이션을 위한 통합적이고 전이 가능한 백본임을 확립한다.
English
Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial priors for visual grounding, spatial reasoning, and pointing, but these capabilities are rarely elicited directly for robot control. Existing navigation systems instead rely on task- or embodiment-specific components, fragmenting perception, reasoning, and action while offering limited generalization. Here we present LightNav-0, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads. LightNav-0 represents diverse navigation tasks through a unified token interface: dual-channel pointing expresses task-, scene-, and embodiment-agnostic spatial intent, while a residual vector-quantized action tokenizer maps this intent to precise, embodiment-specific trajectories. Together with temporally aware visual history compression, ER mid-training, supervised fine-tuning, and reinforcement learning, this formulation supports instruction following, open-vocabulary object navigation, and visual tracking within a single model. The navigation training corpus spans 2K+ scenes and 4K+ hours of embodied navigation data. LightNav-ER, the embodied-reasoning checkpoint used to initialize LightNav-0, attains the highest complete-set average across 8 embodied-reasoning benchmarks, while LightNav-0 achieves state-of-the-art monocular success rates across all 10 public navigation simulation settings. Real-world evaluations further demonstrate zero-shot generalization across robot embodiments, diverse scenes, and static and dynamic targets. These results establish compact VLMs as a unified and transferable backbone for generalist embodied navigation.