遮罩強制:透過雙雜訊遮罩展開改善自迴歸視訊擴散蒸餾
Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
September 8, 2026
作者: Zhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong, Songchun Zhang, Junchao Huang, Guian Fang, Xin Wang, Pan Hui, Anyi Rao
cs.AI
摘要
自迴歸(AR)影片擴散模型在即時影片生成方面展現出巨大潛力。近期方法透過分佈匹配蒸餾(DMD)將預訓練的雙向影片擴散模型蒸餾為因果式 AR 學生模型,但生成的影片常出現過飽和與過平滑問題,導致視覺品質與真實感有限。關鍵成因在於 DMD 中反向 KL 目標的模式尋求行為,可能使學生分佈崩塌到教師分佈的少數模式上。為此,我們提出 Mask Forcing,一種雙雜訊遮罩 Rollout 策略,透過擾動 AR 學生的自我 Rollout 來緩解反向 KL 模式尋求所導致的模式崩塌。其核心思想是在 AR 擴散蒸餾的自我 Rollout 過程中,沿空間與時間軸使用隨機遮罩,將較乾淨的訊號注入雜訊較多的 Rollout 輸入。此類擾動促使學生 Rollout 探索教師分佈的更多區域,使 DMD 能提供學生已覆蓋模式之外的學習訊號。此外,較乾淨的 token 可作為其他雜訊較多 token 的去雜訊引導,改善中間 Rollout 預測並減少誤差累積。大量實驗表明,我們的方法能有效率地提升多種 AR 影片擴散蒸餾方法,並獲得更高的視覺品質,且無需納入真實影片資料或額外後訓練階段。
English
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smoothing issues, resulting in limited visual quality and realism. The key contributing factor is the mode-seeking behavior of the reverse KL objective in DMD, which can cause the student distribution to collapse onto only a few modes of the teacher distribution. To address this, we propose Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking. The core idea is to inject cleaner signals into noisy rollout inputs via random masks along spatial and temporal axes during the self-rollout process of AR diffusion distillation. Such perturbations encourage the student rollouts to explore more regions of the teacher distribution, allowing DMD to provide learning signals beyond the modes already covered by the student. Moreover, the cleaner tokens act as denoising guidance for other noisier tokens, improving the intermediate rollout predictions and reducing error accumulation. Extensive experiments demonstrate that our method improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.