ChatPaper.aiChatPaper

DistillAlign: 자기회귀적 비디오 증류에서 모드 커버링과 모드 시킹의 조정

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

July 29, 2026
저자: Jiaxing Li, Kai Zou, Cindy Zhou, Kaichen Huang, Junyao Gao, Zile Wang, Yang Liu, Bin Liu, Bo An, Yangguang Li
cs.AI

초록

기존의 자기회귀 영상 증류 방법은 일반적으로 분포 정합 증류(DMD) 기반의 다단계 파이프라인을 채택한다. 그러나 이러한 방법들은 초기화 단계와 DMD 단계를 분리하여 각기 다른 목표 분포를 추구하도록 하고, 중간 학생 모델을 주로 VBench와 같은 시각적 점수로 평가하는 경향이 있다. 본 논문에서는 분포 관점에서 이 설계를 재검토한다. 분포 정합 손실의 모드 탐색적 특성을 고려할 때, 좋은 초기화는 단순히 높은 품질을 추구하는 것이 아니라 대상 DMD 교사의 모드 적용 범위와 일치해야 한다. 이를 분석하기 위해, 공유 잠재 공간에서 학생과 교사 분포 간의 정밀도와 적용 범위를 측정하는 분포 평가 프로토콜을 도입한다. 이는 시각적 점수에 가려진 차이를 드러낸다. 일부 초기화는 높은 정밀도에 도달하지만 적용 범위가 낮아 최적의 정제가 이루어지지 않는 반면, 모드를 포괄하는 초기화는 더 넓은 지지를 유지한다. 더 나아가, 목표 분포가 정렬된 경우에도 DMD의 역-KL 목적 함수는 학습 후반에 학생을 높은 확률의 교사 영역으로 유도하여 적용 범위와 다양성을 감소시킬 수 있다. 이 문제를 해결하기 위해, DMD의 모드 탐색 목적과 일관성 증류 기반의 모드 포괄 제약을 결합한 공동 증류를 제안한다. 실험 결과, 본 방법은 생성 품질, 적용 범위, 다양성을 향상시킨다. 특히 Wan-1.3B DMD 교사를 사용하더라도 Wan-14B로 정제된 기준선보다 성능이 뛰어나며, 이는 자기회귀 영상 증류에서 분포 정렬의 중요성을 강조한다.
English
Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they typically decouple the initialization and DMD stages -- which then pursue different target distributions -- and judge the intermediate student mainly by visual scores such as VBench. In this paper, we revisit this design from a distributional perspective. Given the mode-seeking nature of the distribution matching loss, a good initialization should match the mode coverage of the target DMD teacher, rather than merely pursuing high quality. To analyze this, we introduce a distributional evaluation protocol that measures precision and coverage between student and teacher distributions in a shared latent space. It exposes differences hidden by visual scores: some initializations reach high precision but low coverage, leading to suboptimal refinement, while mode-covering ones preserve broader support. Furthermore, even when the target distributions are aligned, DMD's reverse-KL objective can still drive the student toward high-probability teacher regions in late training, reducing coverage and diversity. To address this, we propose joint distillation, which combines DMD's mode-seeking objective with a Consistency Distillation-based mode-covering constraint. Experiments show that our method improves generation quality, coverage, and diversity; notably, even with a Wan-1.3B DMD teacher, it outperforms baselines refined with Wan-14B, underscoring the importance of distributional alignment in autoregressive video distillation.