ChatPaper.aiChatPaper

ワンダー:より良い映像世界モデル

Wonder: Video World Model Done Better

July 28, 2026
著者: Jiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli, Vishal M. Patel, Yiqun Mei
cs.AI

要旨

我々は、リアルタイムかつカメラ制御可能なワールド探索のための汎用ビデオワールドモデル「Wonder」を提案する。画像または条件付きビデオを入力として、Wonderはプレイ可能なワールドを構築する。このワールドでは、ユーザがカメラを動かしてインタラクティブにナビゲートし、未観測領域を発見し、過去に観測した領域をリアルタイムかつ長期にわたって再訪することができる。この能力を実現するには、制御手法、記憶機構、訓練戦略のシステムレベルの協調設計が必要である。本研究では、新しいカメラ条件付け手法として、密な座標フィールドを導入する。そのレンダリングにより、空間的に整列した動きと方向の手がかりが提供され、モデルがカメラの動作を直接視覚的な証拠として解釈できるようになる。また、生成コンテキストが拡大する中での高速かつ正確な記憶検索を実現するため、効率的なスパースアテンションに基づく記憶機構を提案する。これにより、モデルは実際のコンテキスト長にかかわらず、推論時に関連する少数のコンテキストトークンに選択的に注意を向けることができる。さらに、自己強制型蒸留パイプラインを修正するためのいくつかの手法を開発し、生徒モデルが制御信号を尊重する能力を向上させるとともに、教師からの多様な生成モードと長期記憶を維持する。これらの要素を組み合わせることで、Wonderは長期ロールアウトにわたって一貫した幾何、外観、ダイナミクスを保ちながら、16 FPSで多様な分単位のビデオを合成することが可能となる。画像からビデオへの生成に加えて、Wonderはビデオ条件付き生成を自然にサポートし、既存の動的シーンをリアルタイムで再撮影することを可能にする。
English
We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon. Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy. We introduce a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence. To support fast and precise memory retrieval over a growing generation context, we propose an efficient sparse attention-based memory mechanism, enabling the model to selectively attend to a small set of relevant context tokens at inference time, regardless of actual context length. We further develop several techniques to rectify the self-forcing-style distillation pipeline, improving the student model's ability to respect control signals, as well as maintaining diverse generation modes and long-term memory from the teacher. Together, these components enable Wonder to synthesize diverse, minute-scale videos at 16 FPS while preserving coherent geometry, appearance, and dynamics across long rollouts. Beyond image-to-video generation, Wonder naturally supports video-conditioned generation, allowing existing dynamic scenes to be re-shot in real time.