マスク・フォーシング:デュアルノイズマスキング・ロールアウトによる自己回帰型ビデオ拡散蒸留の改善
Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
September 8, 2026
著者: Zhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong, Songchun Zhang, Junchao Huang, Guian Fang, Xin Wang, Pan Hui, Anyi Rao
cs.AI
要旨
自己回帰(AR)動画拡散モデルは、リアルタイム動画生成において大きな可能性を示している。最近の手法では、事前学習済みの双方向動画拡散モデルを分布マッチング蒸留(DMD)によって因果的AR生徒モデルへ蒸留するが、生成される動画は過飽和や過平滑化の問題をしばしば抱え、視覚品質とリアリズムが限られる。主な要因はDMDにおける逆KL目的関数のモード探索挙動であり、これにより生徒分布が教師分布の少数のモードのみに崩壊し得る。この問題に対処するため、我々はMask Forcingを提案する。これは、AR生徒の自己ロールアウトを摂動させ、逆KLモード探索に起因するモード崩壊を緩和するDual-Noise Masking Rollout戦略である。中心となるアイデアは、AR拡散蒸留の自己ロールアウト過程において、空間軸および時間軸に沿ったランダムマスクを介して、よりクリーンな信号をノイズの多いロールアウト入力に注入することである。このような摂動は、生徒のロールアウトが教師分布のより多くの領域を探索することを促し、DMDが生徒によってすでに覆われているモードを超えた学習信号を提供できるようにする。さらに、よりクリーンなトークンは他のよりノイズの多いトークンに対するノイズ除去ガイダンスとして機能し、中間ロールアウト予測を改善し、誤差蓄積を低減する。広範な実験により、我々の手法は実動画データや追加の事後学習段階を組み込むことなく、複数のAR動画拡散蒸留手法をより高い視覚品質で効率的に改善することが示された。
English
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smoothing issues, resulting in limited visual quality and realism. The key contributing factor is the mode-seeking behavior of the reverse KL objective in DMD, which can cause the student distribution to collapse onto only a few modes of the teacher distribution. To address this, we propose Mask Forcing, a Dual-Noise Masking Rollout strategy that perturbs the AR student self-rollout to mitigate mode collapse induced by reverse-KL mode seeking. The core idea is to inject cleaner signals into noisy rollout inputs via random masks along spatial and temporal axes during the self-rollout process of AR diffusion distillation. Such perturbations encourage the student rollouts to explore more regions of the teacher distribution, allowing DMD to provide learning signals beyond the modes already covered by the student. Moreover, the cleaner tokens act as denoising guidance for other noisier tokens, improving the intermediate rollout predictions and reducing error accumulation. Extensive experiments demonstrate that our method improves multiple AR video diffusion distillation methods with higher visual quality efficiently, without incorporating real video data or additional post-training stages.