WorldRover:面向世界探索的可扩展合成视频数据引擎,具有丰富的标注
WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
August 16, 2026
作者: Xiaojie Xu, Zhengyuan Lin, Runyi Li, Yihao Liu, Kaipeng Zhang, Yongtao Ge
cs.AI
摘要
学习生成或重建可探索世界,需要视频不仅配以RGB,还需配以相机运动、场景几何、时间对应关系,以及对于交互模型而言的控制信号。真实采集可以提供部分此类信号,但稠密几何和长程对应通常依赖于估计或专用仪器。渲染能直接提供这些量,但现有的合成资源很少在同一帧上同时包含它们,并支持对视角和外观的受控变化。我们提出WorldRover,这是一个数据引擎,用于生成艺术家构建环境中带有丰富标注的长程探索。其核心WorldRover-Engine是一个Unreal Engine管线,负责执行并离线渲染分钟级路线,同时保留完整的轨迹和场景几何。同一次探索可从第一人称、第三人称和360度全景相机在不同环境状态下重放。利用WorldRover-Engine,我们构建了WorldRover-10M,其序列将RGB与度量深度、相机轨迹以及每次探索中由轨迹导出的动作信号配对。第三人称子集还额外提供稠密光流、带可见性的长程2D/3D点轨迹,以及与相机轨迹不同的角色轨迹。该引擎可以从第一人称、第三人称和360度全景视角渲染遍历过程,支持不同环境状态或使用中性白色材质,同时保持路线和场景几何不变。因此,WorldRover将长时程世界探索转化为一个可扩展的数据生成问题,为那些必须构建、维护并重新访问可探索世界的一致表征的模型提供监督。
English
Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene geometry, temporal correspondence and, for interactive models, control signals. Real capture can provide some of these signals, but dense geometry and long-range correspondence usually rely on estimation or specialised instrumentation. Rendering provides these quantities directly, yet existing synthetic resources rarely combine them on the same frames while also supporting controlled changes of viewpoint and appearance. We introduce WorldRover, a data engine for generating richly annotated, long-range explorations of artist-built environments. At its core, WorldRover-Engine is an Unreal Engine pipeline that executes and offline-renders minute-scale routes while preserving their full trajectories and scene geometry. The same exploration can be replayed from first-person, third-person, and 360 panoramic cameras under different environmental states. Using WorldRover-Engine, we construct WorldRover-10M, whose sequences pair RGB with metric depth, camera trajectories, and trajectory-derived action signals throughout each exploration. Third-person subsets additionally provide dense optical flow, long-range 2D/3D point tracks with visibility, and a character trajectory distinct from the camera trajectory. The engine can render a traversal from first-person, third-person and 360 panoramic viewpoints, under different environmental states or with a neutral white material, while preserving the route and scene geometry. WorldRover therefore turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.