MV-Forcing:基于4D时空自强制的长多视角视频生成
MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing
July 6, 2026
作者: Gal Fiebelman, Hadar Averbuch-Elor, Sagie Benaim
cs.AI
摘要
近期,视频扩散模型的进展使得通过时间自回归生成长时单视角视频,或通过双向注意力生成短时多视角视频成为可能。然而,针对动态场景生成时长一致、多视角连贯的视频仍未得到解决。本文提出MV-Forcing框架,该框架通过在逐次生成的视角之间引入4D几何桥接,将时间与视角维度的自回归整合于单一扩散模型之中。我们的核心洞察在于:自回归3D重建模型能够自然地衔接自回归生成的视角。给定一个已完成的源视角,我们重建其3D结构,并渲染出下一个目标视角的几何先验,扩散模型再将其精化为高质量视频。为了突破教师模型固定时间窗口的限制,我们引入联合去噪机制——训练时两个视角槽均从噪声初始化,从而实现时间无界生成。我们通过时空自强制分布匹配蒸馏(Distribution Matching Distillation with Spatio-Temporal Self-Forcing)对模型进行精炼,弥补了时间与视角序列自回归中训练-推理暴露偏差的差距。在合成数据与真实数据上的大量实验表明,MV-Forcing能够利用单一少步学生模型,以任意时长与视角数量生成动态场景的几何一致多视角视频。
English
Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, generating long, multi-view consistent videos of dynamic scenes remains unsolved. In this work, we present MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. Our key insight is that an autoregressive 3D reconstruction model naturally interfaces between autoregressively generated views. Given a completed source view, we reconstruct its 3D structure and render a geometric prior of the next target viewpoint, which the diffusion model refines into a high-quality video. To extend generation beyond the teacher's fixed temporal window, we introduce a joint denoising regime where both view slots are initialized from noise during training, enabling temporally unbounded generation. We distill the model via Distribution Matching Distillation with Spatio-Temporal Self-Forcing, closing the train-inference exposure bias gap for both temporal and view-sequential autoregression. Extensive experiments on both synthetic and real-world data demonstrate that MV-Forcing produces geometrically consistent multi-view videos of dynamic scenes at arbitrary lengths and viewpoint counts using a single few-step student model.