WorldRover: 풍부한 주석을 제공하는 확장 가능한 합성 비디오 데이터 엔진
WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
August 16, 2026
저자: Xiaojie Xu, Zhengyuan Lin, Runyi Li, Yihao Liu, Kaipeng Zhang, Yongtao Ge
cs.AI
초록
탐험 가능한 세계를 생성하거나 재구성하는 학습은 RGB에 더하여 카메라 모션, 장면 기하 구조, 시간적 대응 관계, 그리고 상호작용 모델의 경우 제어 신호가 함께 제공되는 비디오를 필요로 한다. 실제 촬영은 이러한 신호 중 일부를 제공할 수 있지만, 조밀한 기하 구조와 장기간 대응 관계는 일반적으로 추정이나 특수 계측 장비에 의존한다. 렌더링은 이러한 정보를 직접 제공하지만, 기존의 합성 데이터 자원은 동일한 프레임에 이들을 결합하면서도 시점과 외관의 제어된 변경을 지원하는 경우가 드물다. 아티스트가 제작한 환경에 대한 풍부한 주석이 포함된 장기간 탐험을 생성하는 데이터 엔진인 WorldRover를 소개한다. WorldRover-Engine의 핵심은 언리얼 엔진 파이프라인으로, 분 단위 경로를 실행하여 오프라인으로 렌더링하면서 전체 궤적과 장면 기하 구조를 보존한다. 동일한 탐험은 서로 다른 환경 상태에서 1인칭, 3인칭, 360도 파노라마 카메라로 재생할 수 있다. WorldRover-Engine을 사용하여 WorldRover-10M을 구축하였으며, 이 데이터셋의 시퀀스는 각 탐험 전반에 걸쳐 RGB와 함께 미터 단위 깊이, 카메라 궤적, 궤적에서 파생된 행동 신호를 제공한다. 3인칭 하위 집합은 추가로 조밀한 광류, 가시성 정보를 포함한 장기간 2D/3D 포인트 트랙, 그리고 카메라 궤적과 구별되는 캐릭터 궤적을 제공한다. 엔진은 경로와 장면 기하 구조를 보존하면서, 서로 다른 환경 상태 또는 중립 흰색 재질을 적용한 조건에서 1인칭, 3인칭, 360도 파노라마 시점으로 이동 과정을 렌더링할 수 있다. 따라서 WorldRover는 장기적 세계 탐험을 확장 가능한 데이터 생성 문제로 전환하며, 탐험 가능한 세계에 대한 일관된 표현을 구축·유지·재방문해야 하는 모델에 지도 신호를 제공한다.
English
Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models, control signals. Real capture can provide some of these signals, but dense geometry and long-range correspondence usually rely on estimation or specialised instrumentation. Rendering provides these quantities directly, yet existing synthetic resources rarely combine them on the same frames while also supporting controlled changes of viewpoint and appearance. We introduce WorldRover, a data engine for generating richly annotated, long-range explorations of artist-built environments. At its core, WorldRover-Engine is an Unreal Engine pipeline that executes and offline-renders minute-scale routes while preserving their full trajectories and scene geometry. The same exploration can be replayed from first-person, third-person, and 360 panoramic cameras under different environmental states. Using WorldRover-Engine, we construct WorldRover-10M, whose sequences pair RGB with metric depth, camera trajectories, and trajectory-derived action signals throughout each exploration. Third-person subsets additionally provide dense optical flow, long-range 2D/3D point tracks with visibility, and a character trajectory distinct from the camera trajectory. The engine can render a traversal from first-person, third-person and 360 panoramic viewpoints, under different environmental states or with a neutral white material, while preserving the route and scene geometry. WorldRover therefore turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.