驚奇:更佳影片世界模型
Wonder: Video World Model Done Better
July 28, 2026
作者: Jiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli, Vishal M. Patel, Yiqun Mei
cs.AI
摘要
我們提出 Wonder,一個通用的影片世界模型,用於即時、可控制相機的世界探索。給定一張影像或一段條件影片,Wonder 建構一個可遊玩的世界,使用者可透過移動相機進行互動式導航,即時且長期地探索未見區域,並重新造訪先前觀察過的地方。實現此能力需要系統層級的控制方法、記憶機制與訓練策略的協同設計。我們提出一種新穎的相機條件控制方法,利用密集座標場的渲染結果提供空間對齊的運動與方向線索,使模型能將相機運動直接解讀為視覺證據。為支援在不斷增長的生成上下文中的快速且精確的記憶檢索,我們提出一種基於稀疏注意力機制的有效記憶方法,使模型在推理時能根據實際上下文長度,選擇性地關注少量相關的上下文標記。我們進一步開發多項技術來修正自強迫式蒸餾流程,提升學生模型遵循控制訊號的能力,並維持教師模型的多樣生成模式與長期記憶。這些元件共同使 Wonder 能以每秒 16 幀的速度合成多樣、長達數分鐘的影片,同時在長時間推演中保持一致的幾何、外觀與動態。除了影像到影片的生成,Wonder 也自然支援影片條件生成,允許即時重新拍攝現有的動態場景。
English
We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon. Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy. We introduce a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence. To support fast and precise memory retrieval over a growing generation context, we propose an efficient sparse attention-based memory mechanism, enabling the model to selectively attend to a small set of relevant context tokens at inference time, regardless of actual context length. We further develop several techniques to rectify the self-forcing-style distillation pipeline, improving the student model's ability to respect control signals, as well as maintaining diverse generation modes and long-term memory from the teacher. Together, these components enable Wonder to synthesize diverse, minute-scale videos at 16 FPS while preserving coherent geometry, appearance, and dynamics across long rollouts. Beyond image-to-video generation, Wonder naturally supports video-conditioned generation, allowing existing dynamic scenes to be re-shot in real time.