ChatPaper.aiChatPaper

MV-Forcing: 4D 기반 시공간적 자기 강제를 통한 긴 다중 시점 비디오 생성

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

July 6, 2026
저자: Gal Fiebelman, Hadar Averbuch-Elor, Sagie Benaim
cs.AI

초록

최근 비디오 확산 모델의 발전으로 시간적 자기회귀를 통한 긴 단일 시점 생성 또는 양방향 어텐션을 통한 짧은 다중 시점 합성이 가능해졌다. 그러나 동적 장면의 길고 다중 시점에서 일관된 비디오를 생성하는 문제는 아직 해결되지 않았다. 본 연구에서는 순차적으로 생성된 시점들 사이에 4D 기하학적 브리지를 도입하여 단일 확산 모델 내에서 시간적 및 시점별 자기회귀를 구성하는 프레임워크인 MV-Forcing을 제시한다. 우리의 핵심 통찰은 자기회귀적 3D 복원 모델이 자기회귀적으로 생성된 시점들 사이를 자연스럽게 연결한다는 점이다. 완성된 원본 시점이 주어지면, 우리는 해당 시점의 3D 구조를 복원하고 다음 목표 시점의 기하학적 사전 정보를 렌더링하며, 확산 모델이 이를 고품질 비디오로 정제한다. 교사 모델의 고정된 시간 창을 넘어 생성을 확장하기 위해, 우리는 훈련 중 두 시점 슬롯 모두를 잡음으로 초기화하여 시간적으로 무제한적인 생성을 가능하게 하는 공동 잡음 제거 방식을 도입한다. 우리는 시공간 자기 강제를 적용한 분포 정합 증류를 통해 모델을 증류하여, 시간적 및 시점별 순차 자기회귀 모두에 대한 훈련-추론 노출 편향 차이를 해소한다. 합성 데이터와 실제 데이터 모두에 대한 광범위한 실험을 통해 MV-Forcing이 단일 소수 단계 학생 모델을 사용하여 임의의 길이와 시점 수에서 동적 장면의 기하학적으로 일관된 다중 시점 비디오를 생성함을 입증한다.
English
Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, generating long, multi-view consistent videos of dynamic scenes remains unsolved. In this work, we present MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. Our key insight is that an autoregressive 3D reconstruction model naturally interfaces between autoregressively generated views. Given a completed source view, we reconstruct its 3D structure and render a geometric prior of the next target viewpoint, which the diffusion model refines into a high-quality video. To extend generation beyond the teacher's fixed temporal window, we introduce a joint denoising regime where both view slots are initialized from noise during training, enabling temporally unbounded generation. We distill the model via Distribution Matching Distillation with Spatio-Temporal Self-Forcing, closing the train-inference exposure bias gap for both temporal and view-sequential autoregression. Extensive experiments on both synthetic and real-world data demonstrate that MV-Forcing produces geometrically consistent multi-view videos of dynamic scenes at arbitrary lengths and viewpoint counts using a single few-step student model.