ChatPaper.aiChatPaper

Robostral Navigate

Robostral Navigate

July 22, 2026
저자: Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi, Chris Bamford, Elliot Chane-Sane, Guillaume Lample, Khyathi Raghavi Chandu, Ludovic Ho Fuh, Mathieu Poiree, Olivier Duchenne, Rosalie Millner, Srijan Mishra, Theo Cachet, Thomas Chabal
cs.AI

초록

대규모 내비게이션 시스템을 배포하려면 센서 가정을 최소화하고, 다양한 로봇 형태에 일반화되며, 효율적으로 훈련할 수 있는 레시피가 필요합니다. 그러나 현재 최고의 시스템들은 깊이 센서, 다중 카메라 장치, 또는 사전 구축된 맵에 의존하여 지원하는 하드웨어를 제한하고 배포 비용을 증가시킵니다. 우리는 이러한 확장성 목표를 중심으로 구축된 8B 비전-언어 모델인 Robostral Navigate를 소개합니다. 이 모델은 로봇 플랫폼에서 가장 보편적인 센서인 단일 RGB 이미지 스트림만을 입력으로 사용하며, 현재 카메라 뷰에서 다음 목표 위치를 가리켜 경유점을 예측합니다. 로봇 특정 좌표가 아닌 순수하게 이미지 공간에서 작동함으로써, 정책이 카메라 내부 파라미터와 장면 규모 변화에 자연스럽게 강건해지며, 재보정 없이 바퀴형, 보행형, 항공형 로봇에 배포할 수 있습니다. 실제 데이터 수집에 대한 의존도를 줄이고 쉽게 확장하기 위해 35만 개의 시뮬레이션 장면에서 240만 개의 궤적을 생성합니다. 또한 전체 에피소드를 단일 훈련 시퀀스로 압축하는 프리픽스 캐싱 훈련 레시피를 도입하여 훈련 토큰을 22배 줄이고 훈련 시간을 수개월에서 수일로 단축합니다. 트리 기반 어텐션 마스크는 이전 정답 행동에 대한 조건화를 방지하여 시각적 기반 행동 예측을 장려하며, 강화 학습을 사용하여 탐색 및 회복 능력을 더욱 향상시킵니다. 연속 환경에서의 Room-to-Room 및 Room-Across-Room (R2R-CE 및 RxR-CE) 벤치마크에서 Robostral Navigate는 새로운 최첨단 성능을 설정합니다. R2R-CE에서 단일 RGB 카메라만을 사용하면서도 77.4%의 성공률을 달성하여 최고의 단일 카메라 방식보다 10.5포인트, 가장 강력한 깊이 또는 다중 카메라 시스템보다 5.3포인트 더 높은 성능을 보입니다. RxR-CE에서는 75.1%의 성공률에 도달하여 모든 단일 카메라 기준선을 능가합니다.
English
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduce Robostral Navigate, an 8B vision-language model built around this scalability objective. The model consumes only a stream of monocular RGB images - the most ubiquitous sensor across robotic platforms and predicts waypoints by pointing to the next target location in the current camera view. Operating purely in image space, rather than robot-specific coordinates, makes the policy naturally robust to changes in camera intrinsics and scene scale, enabling deployment across wheeled, legged, and aerial robots without recalibration. We generate 2.4 million trajectories across 350k simulated scenes to reduce the reliance on real-world data collection and scale easily. We further introduce a prefix-caching training recipe that packs entire episodes into single training sequences, reducing training tokens by 22x and cutting training time from months to days. A tree-based attention mask prevents conditioning on previous ground-truth actions, encouraging visually grounded action prediction, and reinforcement learning is used to further improve exploration and recovery capabilities. On the Room-to-Room and Room-Across-Room in Continuous Environments (R2R-CE and RxR-CE) benchmarks, Robostral Navigate sets a new state of the art. On R2R-CE, it achieves a 77.4% success rate, surpassing the best monocular method by 10.5 points and the strongest depth- or multi-camera system by 5.3 points despite using only a single RGB camera. On RxR-CE, it reaches 75.1% success rate, outperforming all monocular baselines.