ChatPaper.aiChatPaper

Image2Sim: 生成型ニューラルシミュレータによる身体性ナビゲーションのスケーリング

Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

July 7, 2026
著者: Zihan Wang, Seungjun Lee, Yinghao Xu, Gim Hee Lee
cs.AI

要旨

身体化ナビゲーションは、マルチモーダルな目標を解釈し、3次元空間で推論し、現実世界で確実に目的地に到達するエージェントを構築することを目的としている。しかし、スケーラブルで忠実度が高く、物理的に基づいたインタラクティブな環境が不足しているため、進展は依然として制限されている。実際のスキャンデータセットは視覚的な現実感を提供するが、スケールに制限がある。一方、合成シミュレータはより簡単にスケールできるが、多くの場合シミュレーションと現実のギャップが大きい。我々は、ポーズ付きRGB-D画像シーケンスから高品質なインタラクティブ環境を構築するリアルタイムニューラルシミュレーションフレームワーク、Image2Simを紹介する。中心となるアイデアは、3次元空間アンカリングとフォトリアリスティックな観測合成を分離することである。シーン構築において、Image2Simはフィードフォワード特徴ガウシアンモデルを使用し、ポーズ付きRGB-D観測を単一パスで3次元特徴ガウシアン表現に変換する。レンダリングについては、スパースでノイズの多いガウシアン投影を高品質なパノラマRGB-D観測に変換するGeometry-Aware One-Step Pixel Flowモデルを提案する。Image2Simはまた、高忠実度の観測、実行可能なアクション、多様なナビゲーション指示を大規模に生成する完全自動化された身体化データエンジンとして機能する。大規模なビデオ・画像コレクションを約2万のインタラクティブシーンに変換し、1000万以上のナビゲーション訓練サンプルを合成する。これらのニューラル環境で完全に訓練されたナビゲーションモデルは、主要なベンチマークで大きな改善を達成し、現実世界のゼロショット設定にも効果的に転移する。これらの結果は、スケーラブルなニューラルシミュレーションが、大規模な身体化ナビゲーションの実用的な訓練基盤として機能する可能性を示唆している。
English
Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lack of scalable, high-fidelity, and physically grounded interactive environments. Although real-world scanned datasets offer visual realism, they are limited by scale. In contrast, synthetic simulators scale more easily but often exhibit large sim-to-real gaps. We introduce Image2Sim, a real-time neural simulation framework that constructs high-quality interactive environments from posed RGB-D image sequences. The central idea is to decouple 3D spatial anchoring from photorealistic observation synthesis. For scene construction, Image2Sim uses a feed-forward feature Gaussian model that lifts posed RGB-D observations into a 3D feature-Gaussian representation in a single pass. For rendering, we propose a Geometry-Aware One-Step Pixel Flow model that transforms sparse and noisy Gaussian projections into high-quality panoramic RGB-D observations. Image2Sim also serves as a fully automated embodied data engine that generates high-fidelity observations, executable actions, and diverse navigation instructions at scale. It converts large collections of videos and images into nearly 20K interactive scenes and synthesizes more than 10 million navigation training samples. Navigation models trained entirely in these neural environments achieve strong improvements on major benchmarks and transfer effectively to real-world zero-shot settings. These results suggest that scalable neural simulation can serve as a practical training substrate for embodied navigation at scale.