ChatPaper.aiChatPaper

MV-Forcing: 4Dに基づく時空間自己強制による長尺多視点ビデオ生成

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

July 6, 2026
著者: Gal Fiebelman, Hadar Averbuch-Elor, Sagie Benaim
cs.AI

要旨

近年のビデオ拡散モデルの進歩により、時間的自己回帰による長い単視点生成、または双方向注意機構による短い多視点合成が可能になりました。しかし、動的シーンの長く多視点で一貫性のあるビデオを生成することは未解決の課題です。本研究では、逐次的に生成される視点間に4次元的な幾何学的ブリッジを導入することで、単一の拡散モデル内で時間的および視点方向の自己回帰を組み合わせるフレームワーク、MV-Forcingを提案します。我々の重要な洞察は、自己回帰的3D再構成モデルが自己回帰的に生成された視点間を自然に仲介するという点です。完了したソース視点が与えられると、その3D構造を再構成し、次のターゲット視点の幾何学的事前情報をレンダリングします。これを拡散モデルが高品質なビデオに洗練します。教師モデルの固定された時間窓を超えて生成を拡張するために、訓練中に両方の視点スロットをノイズから初期化する共同ノイズ除去方式を導入し、時間的に制限のない生成を可能にします。我々は、時空間的自己強制(Spatio-Temporal Self-Forcing)を伴う分布マッチング蒸留(Distribution Matching Distillation)によりモデルを蒸留し、時間的および視点順序的自己回帰の両方について訓練-推論の露出バイアスギャップを解消します。合成データと実世界データの両方での広範な実験により、MV-Forcingが単一の数ステップの生徒モデルを使用して、任意の長さと視点数の動的シーンの幾何学的に一貫した多視点ビデオを生成できることを示します。
English
Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, generating long, multi-view consistent videos of dynamic scenes remains unsolved. In this work, we present MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. Our key insight is that an autoregressive 3D reconstruction model naturally interfaces between autoregressively generated views. Given a completed source view, we reconstruct its 3D structure and render a geometric prior of the next target viewpoint, which the diffusion model refines into a high-quality video. To extend generation beyond the teacher's fixed temporal window, we introduce a joint denoising regime where both view slots are initialized from noise during training, enabling temporally unbounded generation. We distill the model via Distribution Matching Distillation with Spatio-Temporal Self-Forcing, closing the train-inference exposure bias gap for both temporal and view-sequential autoregression. Extensive experiments on both synthetic and real-world data demonstrate that MV-Forcing produces geometrically consistent multi-view videos of dynamic scenes at arbitrary lengths and viewpoint counts using a single few-step student model.