Mask Forcing: 이중 노이즈 마스킹 롤아웃을 통한 자기회귀 비디오 확산 증류 개선
Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
September 8, 2026
저자: Zhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong, Songchun Zhang, Junchao Huang, Guian Fang, Xin Wang, Pan Hui, Anyi Rao
cs.AI
초록
자기회귀(AR) 비디오 확산 모델은 실시간 비디오 생성에서 큰 잠재력을 보여 왔다. 최근 방법들은 사전 학습된 양방향 비디오 확산 모델을 분포 매칭 증류(DMD)를 통해 인과적 AR 학생 모델로 증류하지만, 생성된 비디오는 종종 과포화 및 과평활화 문제를 겪어 제한된 시각적 품질과 사실감을 보인다. 주요 원인은 DMD에서 역 KL 목적 함수의 모드 추구(mode-seeking) 동작으로, 이는 학생 분포가 교사 분포의 소수 모드에만 붕괴되도록 할 수 있다. 이를 해결하기 위해, 우리는 AR 학생의 자기 롤아웃을 교란하여 역-KL 모드 추구로 인한 모드 붕괴를 완화하는 이중 노이즈 마스킹 롤아웃(Dual-Noise Masking Rollout) 전략인 Mask Forcing을 제안한다. 핵심 아이디어는 AR 확산 증류의 자기 롤아웃 과정에서 공간 및 시간 축을 따른 무작위 마스크를 통해 잡음이 있는 롤아웃 입력에 더 깨끗한 신호를 주입하는 것이다. 이러한 교란은 학생 롤아웃이 교사 분포의 더 많은 영역을 탐색하도록 장려하여, DMD가 학생이 이미 커버한 모드 너머의 학습 신호를 제공할 수 있게 한다. 게다가, 더 깨끗한 토큰은 다른 더 잡음이 많은 토큰에 대한 노이즈 제거 지침으로 작용하여 중간 롤아웃 예측을 개선하고 오류 누적을 줄인다. 광범위한 실험은 우리 방법이 실제 비디오 데이터나 추가적인 후학습 단계를 포함하지 않고도 여러 AR 비디오 확산 증류 방법을 더 높은 시각적 품질로 효율적으로 개선함을 입증한다.
English
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smoothing issues, resulting in limited visual quality and realism. The key contributing factor is the mode-seeking behavior of the reverse KL objective in DMD, which can cause the student distribution to collapse onto only a few modes of the teacher distribution. To address this, we propose Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking. The core idea is to inject cleaner signals into noisy rollout inputs via random masks along spatial and temporal axes during the self-rollout process of AR diffusion distillation. Such perturbations encourage the student rollouts to explore more regions of the teacher distribution, allowing DMD to provide learning signals beyond the modes already covered by the student. Moreover, the cleaner tokens act as denoising guidance for other noisier tokens, improving the intermediate rollout predictions and reducing error accumulation. Extensive experiments demonstrate that our method improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.