ロボストラル・ナビゲート
Robostral Navigate
July 22, 2026
著者: Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi, Chris Bamford, Elliot Chane-Sane, Guillaume Lample, Khyathi Raghavi Chandu, Ludovic Ho Fuh, Mathieu Poiree, Olivier Duchenne, Rosalie Millner, Srijan Mishra, Theo Cachet, Thomas Chabal
cs.AI
要旨
大規模なナビゲーションシステムを展開するには、センサーの前提条件を最小限に抑え、ロボットの形態を超えて汎化し、効率的に訓練できる手法が必要です。しかし、現在の最先端システムは深度センサーやマルチカメラリグ、事前構築された地図に依存しており、対応可能なハードウェアが制限され、導入コストが増大しています。本論文では、このスケーラビリティ目標に基づいて構築された8Bパラメータの視覚言語モデル「Robostral Navigate」を紹介します。本モデルは、ロボットプラットフォームで最も広く普及したセンサーである単眼RGB画像のストリームのみを入力として、現在のカメラ視野内の次の目標位置を指示することでウェイポイントを予測します。ロボット固有の座標ではなく、純粋に画像空間で動作することで、カメラ内部パラメータやシーンスケールの変化に対してポリシーが自然にロバストとなり、再較正なしで車輪型、脚型、空中ロボットにわたって展開可能です。現実世界でのデータ収集への依存を減らし、容易にスケールアップするために、35万のシミュレーションシーンで240万の軌跡を生成しました。さらに、エピソード全体を単一の訓練シーケンスにパッケージ化するプレフィックスキャッシング訓練手法を導入し、訓練トークンを22倍削減し、訓練時間を月単位から日単位に短縮しました。ツリー構造のアテンションマスクは過去の正解行動に基づく条件付けを防ぎ、視覚に基づいた行動予測を促進します。また、強化学習を用いて探索能力と回復能力をさらに向上させました。連続環境におけるRoom-to-RoomおよびRoom-Across-Room(R2R-CEおよびRxR-CE)ベンチマークにおいて、Robostral Navigateは新たな最先端を達成しました。R2R-CEでは成功率77.4%を達成し、最良の単眼方式を10.5ポイント上回り、単一RGBカメラのみを使用しながら、深度センサーやマルチカメラシステムを採用した最強の方式をも5.3ポイント凌駕しました。RxR-CEでは成功率75.1%に達し、すべての単眼ベースラインを上回りました。
English
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduce Robostral Navigate, an 8B vision-language model built around this scalability objective. The model consumes only a stream of monocular RGB images - the most ubiquitous sensor across robotic platforms and predicts waypoints by pointing to the next target location in the current camera view. Operating purely in image space, rather than robot-specific coordinates, makes the policy naturally robust to changes in camera intrinsics and scene scale, enabling deployment across wheeled, legged, and aerial robots without recalibration. We generate 2.4 million trajectories across 350k simulated scenes to reduce the reliance on real-world data collection and scale easily. We further introduce a prefix-caching training recipe that packs entire episodes into single training sequences, reducing training tokens by 22x and cutting training time from months to days. A tree-based attention mask prevents conditioning on previous ground-truth actions, encouraging visually grounded action prediction, and reinforcement learning is used to further improve exploration and recovery capabilities. On the Room-to-Room and Room-Across-Room in Continuous Environments (R2R-CE and RxR-CE) benchmarks, Robostral Navigate sets a new state of the art. On R2R-CE, it achieves a 77.4% success rate, surpassing the best monocular method by 10.5 points and the strongest depth- or multi-camera system by 5.3 points despite using only a single RGB camera. On RxR-CE, it reaches 75.1% success rate, outperforming all monocular baselines.