ChatPaper.aiChatPaper

Open-MOPD: 다중 교사 온-정책 증류의 능력 불균형 진단 및 해결

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

August 19, 2026
저자: Huan-ang Gao, Haohan Chi, Yong Yan, Shiyuan Feng, Hanlin Wu, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
cs.AI

초록

다중 교사 온-정책 증류(M-OPD)는 밀집된 토큰 수준 보상 감독을 통해 도메인 특화 강화 학습(RL) 전문가들을 단일 범용 학생 모델로 통합하는 유망한 패러다임으로 부상하였다. 실용적 성공에도 불구하고, 다중 교사 능력 통합을 지배하는 최적화 동역학은 여전히 제대로 이해되지 못하고 있으며, 개방적이고 엄격하게 재현 가능한 레시피가 현저히 부족한 실정이다. 본 연구에서는 오라클 라우팅을 적용한 SmolLM3-3B-Base 기반의 통제된 M-OPD 벤치마크를 구축하여, 능력 통합을 라우팅 모호성으로부터 분리하였다. 우리의 조사는 두드러진 능력 통합 격차를 밝혀내는데, 표준 M-OPD는 도메인 라우팅 오라클 앙상블 대비 가용한 개선 여지의 35.6%만을 포착하며, 지시 따르기와 같은 간결한 작업은 심각한 성능 저하와 조기 수렴 정체를 겪는다. 핵심적으로, 이러한 실패는 그래디언트 충돌이 아닌 토큰 수준 최적화 예산의 심각한 잘못된 배분에서 비롯됨을 우리는 입증한다. 이러한 병폐는 세 가지 직교 요인에 의해 유발된다: 도메인 간 구조적 시퀀스 길이 불일치, 비균일 학습률로 인한 동적 수렴 드리프트, 그리고 비동기 정책 업데이트로 인한 다단계 보상 지연이다. 이러한 불균형을 해소하기 위해 우리는 토큰 비중 균형, 격차 인지 동적 예산 할당, 학생 보상 갱신을 통합하는 원리 기반 프레임워크인 Open-MOPD를 도입한다. 이 메커니즘들은 함께 작동하여 교차 도메인 균형을 체계적으로 회복하며, 단일 배포 가능한 학생 모델에서 개선 여지 회복율을 35.6%에서 83.4%로 끌어올린다. 우리는 학술적으로 접근 가능한 하드웨어 예산 내에서 종단 간 사후 훈련 레시피, 훈련 궤적, 평가 스위트를 완전히 오픈소스로 공개한다.
English
Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learning (RL) experts into a single generalist student via dense, token-level reward supervision. Despite its practical success, the optimization dynamics governing multi-teacher capability integration remain poorly understood, and open, rigorously reproducible recipes are conspicuously lacking. In this work, we establish a controlled M-OPD benchmark on SmolLM3-3B-Base with oracle routing, isolating capability integration from routing ambiguity. Our investigation reveals a pronounced capability integration gap: standard M-OPD captures only 35.6% of the available headroom relative to a domain-routed oracle ensemble, with concise tasks such as instruction following suffering severe degradation and premature stagnation. Crucially, we show that this failure stems not from gradient conflict, but from a severe misallocation of the token-level optimization budget. This pathology is driven by three orthogonal factors: structural sequence-length disparities across domains, dynamic convergence drift due to non-uniform learning rates, and multi-step reward staleness from asynchronous policy updates. To resolve these imbalances, we introduce Open-MOPD, a principled framework incorporating token-share balancing, gap-aware dynamic budget allocation, and student reward refresh. Together, these mechanisms systematically restore cross-domain balance, elevating headroom recovery from 35.6% to 83.4% in a single deployable student. We fully open-source our end-to-end post-training recipe, training trajectories, and evaluation suites on an academically accessible hardware budget.