ChatPaper.aiChatPaper

Wonder:做得更好的视频世界模型

Wonder: Video World Model Done Better

July 28, 2026
作者: Jiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli, Vishal M. Patel, Yiqun Mei
cs.AI

摘要

我们提出Wonder,一种通用视频世界模型,专为实时、摄像机可控的世界探索而设计。给定一张图像或一段条件视频,Wonder构建了一个可交互操作的虚拟世界,用户可通过移动摄像机实时、长期地导航、探索未见区域并重新访问已观察区域。实现这一能力需要系统级协同设计控制方法、记忆机制和训练策略。我们引入一种新颖的摄像机条件控制方法——使用密集坐标场,其渲染结果提供空间对齐的运动与方向线索,使模型能够将摄像机运动直接解释为视觉证据。为支持不断增长的生成上下文中的快速精准记忆检索,我们提出一种基于稀疏注意力的高效记忆机制,使模型在推理时能够选择性关注少量相关上下文标记,而无需考虑实际上下文长度。我们进一步开发多种技术修正自强制式蒸馏流程,提升学生模型遵循控制信号的能力,同时保持教师模型的多样化生成模式与长期记忆。这些组件的结合使Wonder能够以16 FPS速率合成多样化的分钟级视频,并在长程生成中保持几何、外观与动态的连贯性。除图像到视频的生成外,Wonder自然支持视频条件生成,使现有动态场景能够被实时重新拍摄。
English
We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon. Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy. We introduce a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence. To support fast and precise memory retrieval over a growing generation context, we propose an efficient sparse attention-based memory mechanism, enabling the model to selectively attend to a small set of relevant context tokens at inference time, regardless of actual context length. We further develop several techniques to rectify the self-forcing-style distillation pipeline, improving the student model's ability to respect control signals, as well as maintaining diverse generation modes and long-term memory from the teacher. Together, these components enable Wonder to synthesize diverse, minute-scale videos at 16 FPS while preserving coherent geometry, appearance, and dynamics across long rollouts. Beyond image-to-video generation, Wonder naturally supports video-conditioned generation, allowing existing dynamic scenes to be re-shot in real time.