ChatPaper.aiChatPaper

WorldRover:一個用於世界探索、具備豐富標註的可擴展合成影片資料引擎

WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations

August 16, 2026
作者: Xiaojie Xu, Zhengyuan Lin, Runyi Li, Yihao Liu, Kaipeng Zhang, Yongtao Ge
cs.AI

摘要

學習生成或重建可探索的世界,需要將影片與超越 RGB 的更多訊號配對:相機運動、場景幾何、時間對應,以及針對互動模型的控制訊號。真實拍攝能提供其中部分訊號,但稠密幾何與長距離對應通常依賴估計或專門儀器。渲染能直接提供這些量,但現有的合成資源很少在同一幀上將它們結合,同時又支援對視角與外觀的受控變化。我們提出 WorldRover,這是一個資料引擎,用於生成對藝術家建構環境進行豐富標註的長距離探索。其核心 WorldRover-Engine 是一套 Unreal Engine 管線,可執行並離線渲染分鐘級的路徑,同時保留完整的軌跡與場景幾何。相同的探索可從第一人稱、第三人稱及 360 度全景相機,在不同的環境狀態下重播。利用 WorldRover-Engine,我們建構了 WorldRover-10M,其序列在每次探索中將 RGB 與公制深度、相機軌跡,以及由軌跡推導的動作訊號配對。第三人稱子集額外提供稠密光流、具有可見性的長距離 2D/3D 點軌跡,以及不同於相機軌跡的角色軌跡。該引擎能從第一人稱、第三人稱及 360 度全景視點渲染一次遍歷,可在不同的環境狀態下或使用中性白色材質,同時保留路徑與場景幾何。因此,WorldRover 將長程的世界探索轉化為一個可擴展的資料生成問題,為必須建構、維護並重新檢視可探索世界之連貫表徵的模型提供監督。
English
Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models, control signals. Real capture can provide some of these signals, but dense geometry and long-range correspondence usually rely on estimation or specialised instrumentation. Rendering provides these quantities directly, yet existing synthetic resources rarely combine them on the same frames while also supporting controlled changes of viewpoint and appearance. We introduce WorldRover, a data engine for generating richly annotated, long-range explorations of artist-built environments. At its core, WorldRover-Engine is an Unreal Engine pipeline that executes and offline-renders minute-scale routes while preserving their full trajectories and scene geometry. The same exploration can be replayed from first-person, third-person, and 360 panoramic cameras under different environmental states. Using WorldRover-Engine, we construct WorldRover-10M, whose sequences pair RGB with metric depth, camera trajectories, and trajectory-derived action signals throughout each exploration. Third-person subsets additionally provide dense optical flow, long-range 2D/3D point tracks with visibility, and a character trajectory distinct from the camera trajectory. The engine can render a traversal from first-person, third-person and 360 panoramic viewpoints, under different environmental states or with a neutral white material, while preserving the route and scene geometry. WorldRover therefore turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.