WorldRover: リッチなアノテーションを備えた世界探査のためのスケーラブルな合成ビデオデータエンジン
WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
August 16, 2026
著者: Xiaojie Xu, Zhengyuan Lin, Runyi Li, Yihao Liu, Kaipeng Zhang, Yongtao Ge
cs.AI
要旨
探索可能な世界を生成または再構成する学習には、RGBだけでなく、カメラモーション、シーン幾何、時間的対応、そしてインタラクティブモデルについては制御信号と併せたビデオが必要である。実キャプチャはこれらの信号の一部を提供できるが、密な幾何や長距離対応は通常、推定または専用機器に依存する。レンダリングはこれらの量を直接提供するが、既存の合成リソースが同じフレーム上でそれらを組み合わせ、かつ視点と外観の制御された変更をサポートすることはほとんどない。我々は、アーティスト制作環境の豊富な注釈付き長距離探索を生成するためのデータエンジンであるWorldRoverを紹介する。その中核となるWorldRover-Engineは、Unreal Engineパイプラインであり、分単位のルートを実行およびオフラインレンダリングしながら、その完全な軌道とシーン幾何を保持する。同じ探索は、異なる環境状態の下で、一人称、三人称、および360度パノラマカメラから再生することができる。WorldRover-Engineを用いて、我々はWorldRover-10Mを構築する。そのシーケンスは、各探索全体を通じてRGBをメートル単位の深度、カメラ軌道、および軌道由来のアクション信号とペアにする。三人称サブセットはさらに、密なオプティカルフロー、可視性を伴う長距離2D/3Dポイントトラック、およびカメラ軌道とは異なるキャラクター軌道を提供する。このエンジンは、ルートとシーン幾何を保持しながら、異なる環境状態の下で、またはニュートラルな白マテリアルを用いて、一人称、三人称、および360度パノラマ視点からトラバーサルをレンダリングできる。したがってWorldRoverは、長期的な世界探索をスケーラブルなデータ生成問題へと転換し、探索可能な世界の一貫した表現を構築、維持、再訪しなければならないモデルに教師信号を提供する。
English
Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models, control signals. Real capture can provide some of these signals, but dense geometry and long-range correspondence usually rely on estimation or specialised instrumentation. Rendering provides these quantities directly, yet existing synthetic resources rarely combine them on the same frames while also supporting controlled changes of viewpoint and appearance. We introduce WorldRover, a data engine for generating richly annotated, long-range explorations of artist-built environments. At its core, WorldRover-Engine is an Unreal Engine pipeline that executes and offline-renders minute-scale routes while preserving their full trajectories and scene geometry. The same exploration can be replayed from first-person, third-person, and 360 panoramic cameras under different environmental states. Using WorldRover-Engine, we construct WorldRover-10M, whose sequences pair RGB with metric depth, camera trajectories, and trajectory-derived action signals throughout each exploration. Third-person subsets additionally provide dense optical flow, long-range 2D/3D point tracks with visibility, and a character trajectory distinct from the camera trajectory. The engine can render a traversal from first-person, third-person and 360 panoramic viewpoints, under different environmental states or with a neutral white material, while preserving the route and scene geometry. WorldRover therefore turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.