D³-MOPD: 효율적인 멀티-티처 증류를 위한 적응형 동적 도메인 스케줄링
D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation
August 25, 2026
저자: Zechen Sun, Zhiwei Zhang, Fei Zhao, Juntao Li, Mu Chuan, Huayu Deng, Guojian Zhan, Wenliang Chen, Yao Hu, Min Zhang
cs.AI
초록
다중 교사 온-폴리시 증류(MOPD)는 학생 모델 자체 롤아웃에서 도메인별 역방향 KL 발산을 최소화하여 여러 도메인 전문가 교사를 단일 학생 모델로 증류하는 기법이다. 기존 접근법은 일반적으로 훈련 전에 도메인별 데이터 혼합 비율을 고정하지만, 도메인마다 수렴 속도가 크게 다르다는 사실을 간과한다. 일부 도메인은 조기에 정체되는 반면, 다른 도메인은 훈련 예산 전반에 걸쳐 계속 개선된다. 따라서 고정된 혼합 비율은 빠르게 수렴하는 도메인에 계산 자원을 낭비하고, 느리게 수렴하는 도메인은 충분히 훈련하지 못한다. 이러한 문제를 해결하기 위해, 우리는 훈련 중 이미 생성되는 도메인별 역방향 KL 신호를 재활용하여 도메인 혼합 비율을 온라인으로 적응 조정하는 제로 오버헤드 스케줄러인 D³-MOPD(동적 도메인 스케줄링 기반 MOPD)를 제안한다. 이는 훈련 프로세스 외부에서 비동기적으로 실행되며, 프로세스 외부 감시자가 각 도메인의 KL 궤적을 주기적으로 추적하여 남은 개선 여지와 현재 개선 속도를 추정하고, 핵심 훈련 루프를 변경하지 않고 도메인 샘플링 비율을 조정한다. D³-MOPD는 임의의 수의 도메인으로 자연스럽게 확장되며, 더 많은 도메인이 더 다양한 수렴 패턴을 도입할수록 기대 이점이 커진다. 4개의 도메인 전문가 교사로부터 증류된 Qwen3.6-35B-A3B 학생 모델에서 D³-MOPD는 평균 학생-교사 성능 격차의 97%를 좁혀 기본 MOPD의 63%를 크게 상회하며, 약 3배 적은 롤아웃 단계로 동일한 최고 성능에 도달하고, 7개 벤치마크 중 3개에서 전문가 교사를 능가한다.
English
Multi-teacher on-policy distillation (MOPD) distills several domain-expert teachers into a single student by minimizing per-domain reverse-KL divergence on the student's own rollouts. Existing approaches typically fix the per-domain data mixture before training, overlooking the fact that different domains converge at substantially different rates: some plateau early while others continue to improve throughout the training budget. A fixed mixture therefore wastes compute on fast-converging domains and undertrains slower-converging ones. To address this, we propose D^3-MOPD (Dynamic Domain ScheDuling for MOPD), a zero-overhead scheduler that repurposes the per-domain reverse-KL signal already produced during training to adapt the domain mixture online. Running asynchronously outside the training process, an off-process watcher periodically tracks each domain's KL trajectory, estimates remaining headroom and current improvement rate, and accordingly adjusts the domain sampling ratios without altering the core training loop. Our D^3-MOPD scales naturally to arbitrary numbers of domains, and the expected benefit grows as more domains introduce more diverse convergence patterns for the scheduler to exploit. On a Qwen3.6-35B-A3B student distilled from four domain-expert teachers, D^3-MOPD closes 97% of the average student-to-teacher performance gap, compared with 63% for vanilla MOPD, reaches the same peak performance with an approximately 3times reduction in rollout steps, and surpasses the specialist teachers on three of seven benchmarks.