DistillAlign: 自己回帰型ビデオ蒸留におけるモード網羅とモード探索の協調
DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
July 29, 2026
著者: Jiaxing Li, Kai Zou, Cindy Zhou, Kaichen Huang, Junyao Gao, Zile Wang, Yang Liu, Bin Liu, Bo An, Yangguang Li
cs.AI
要旨
既存の自己回帰型ビデオ蒸留手法は、一般的に分布マッチング蒸留(DMD)に基づく多段階パイプラインを採用しています。しかし、これらの手法は初期化段階とDMD段階を分離して扱い、それぞれ異なる目標分布を追求する一方で、中間的な生徒モデルを主にVBenchなどの視覚スコアで評価する傾向があります。本論文では、この設計を分布の観点から再検討します。分布マッチング損失が持つモード探索的な性質を考慮すると、優れた初期化とは単に高品質を追求するのではなく、目標とするDMD教師のモード被覆に適合するものであるべきです。この分析のため、共有潜在空間における生徒分布と教師分布の精度と被覆率を測定する分布的評価プロトコルを導入します。この手法により、視覚スコアでは隠されていた差異が明らかになります。すなわち、一部の初期化は高い精度に達するものの被覆率が低く、最適な精緻化に至らない一方、モード被覆型の初期化はより広いサポートを保持します。さらに、目標分布が一致している場合でも、DMDの逆KL目的関数は学習後期において生徒を教師の高確率領域へと誘導し、被覆率と多様性を低下させる可能性があります。この問題に対処するため、DMDのモード探索目的関数と一致性蒸留に基づくモード被覆制約を組み合わせた、統合蒸留を提案します。実験結果は、本手法が生成品質、被覆率、多様性を改善することを示しています。特筆すべきは、Wan-1.3B DMD教師を用いた場合でも、Wan-14Bで精緻化されたベースラインを上回る性能を達成しており、自己回帰型ビデオ蒸留における分布的一致の重要性を強調しています。
English
Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they typically decouple the initialization and DMD stages -- which then pursue different target distributions -- and judge the intermediate student mainly by visual scores such as VBench. In this paper, we revisit this design from a distributional perspective. Given the mode-seeking nature of the distribution matching loss, a good initialization should match the mode coverage of the target DMD teacher, rather than merely pursuing high quality. To analyze this, we introduce a distributional evaluation protocol that measures precision and coverage between student and teacher distributions in a shared latent space. It exposes differences hidden by visual scores: some initializations reach high precision but low coverage, leading to suboptimal refinement, while mode-covering ones preserve broader support. Furthermore, even when the target distributions are aligned, DMD's reverse-KL objective can still drive the student toward high-probability teacher regions in late training, reducing coverage and diversity. To address this, we propose joint distillation, which combines DMD's mode-seeking objective with a Consistency Distillation-based mode-covering constraint. Experiments show that our method improves generation quality, coverage, and diversity; notably, even with a Wan-1.3B DMD teacher, it outperforms baselines refined with Wan-14B, underscoring the importance of distributional alignment in autoregressive video distillation.