Image2Sim: 생성 신경 시뮬레이터를 통한 체화된 네비게이션 확장
Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator
July 7, 2026
저자: Zihan Wang, Seungjun Lee, Yinghao Xu, Gim Hee Lee
cs.AI
초록
임베디드 내비게이션(embodied navigation)은 다중 모드 목표를 해석하고, 3D 공간에서 추론하며, 실제 세계에서 목표 지점에 안정적으로 도달하는 에이전트를 구축하는 것을 목표로 한다. 그러나 확장 가능하고, 높은 충실도를 가지며, 물리적으로 기반한 상호작용 환경의 부족으로 인해 발전이 제약을 받고 있다. 실제 세계 스캔 데이터셋은 시각적 현실감을 제공하지만 규모 면에서 한계가 있다. 반면, 합성 시뮬레이터는 더 쉽게 확장 가능하지만 종종 큰 시뮬레이션-현실(sim-to-real) 간극을 보인다. 본 논문에서는 포즈가 지정된 RGB-D 이미지 시퀀스로부터 고품질 상호작용 환경을 구축하는 실시간 신경 시뮬레이션 프레임워크인 Image2Sim을 소개한다. 핵심 아이디어는 3D 공간 앵커링을 사실적인 관측 합성으로부터 분리하는 것이다. 장면 구축을 위해 Image2Sim은 피드포워드 특징 가우시안 모델을 사용하여 포즈가 지정된 RGB-D 관측을 단일 패스로 3D 특징-가우시안 표현으로 변환한다. 렌더링을 위해, 우리는 기하학 인식 원스텝 픽셀 플로우(Geometry-Aware One-Step Pixel Flow) 모델을 제안하여 희소하고 잡음이 있는 가우시안 프로젝션을 고품질 파노라마 RGB-D 관측으로 변환한다. Image2Sim은 또한 고충실도 관측, 실행 가능한 행동, 다양한 내비게이션 명령을 대규모로 생성하는 완전 자동화된 임베디드 데이터 엔진 역할을 한다. 대규모 비디오 및 이미지 컬렉션을 거의 20K개의 상호작용 장면으로 변환하고 1천만 개 이상의 내비게이션 훈련 샘플을 합성한다. 이러한 신경 환경에서 전적으로 훈련된 내비게이션 모델은 주요 벤치마크에서 큰 성능 향상을 달성하고 실제 세계 제로샷 설정으로 효과적으로 전이된다. 이러한 결과는 확장 가능한 신경 시뮬레이션이 대규모 임베디드 내비게이션을 위한 실용적인 훈련 기반이 될 수 있음을 시사한다.
English
Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lack of scalable, high-fidelity, and physically grounded interactive environments. Although real-world scanned datasets offer visual realism, they are limited by scale. In contrast, synthetic simulators scale more easily but often exhibit large sim-to-real gaps. We introduce Image2Sim, a real-time neural simulation framework that constructs high-quality interactive environments from posed RGB-D image sequences. The central idea is to decouple 3D spatial anchoring from photorealistic observation synthesis. For scene construction, Image2Sim uses a feed-forward feature Gaussian model that lifts posed RGB-D observations into a 3D feature-Gaussian representation in a single pass. For rendering, we propose a Geometry-Aware One-Step Pixel Flow model that transforms sparse and noisy Gaussian projections into high-quality panoramic RGB-D observations. Image2Sim also serves as a fully automated embodied data engine that generates high-fidelity observations, executable actions, and diverse navigation instructions at scale. It converts large collections of videos and images into nearly 20K interactive scenes and synthesizes more than 10 million navigation training samples. Navigation models trained entirely in these neural environments achieve strong improvements on major benchmarks and transfer effectively to real-world zero-shot settings. These results suggest that scalable neural simulation can serve as a practical training substrate for embodied navigation at scale.