Open-MOPD:診斷與修復多教師同策略蒸餾中的能力不平衡
Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
August 19, 2026
作者: Huan-ang Gao, Haohan Chi, Yong Yan, Shiyuan Feng, Hanlin Wu, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
cs.AI
摘要
多教師在線策略蒸餾(M-OPD)已成為一種前景廣闊的範式,透過密集的詞元級獎勵監督,將領域專精的強化學習專家整合為單一的通用型學生模型。儘管其在實踐上取得了成功,但驅動多教師能力整合的優化動態機制仍未被充分理解,且明顯缺乏公開且可嚴格重現的完整方案。在本研究中,我們在 SmolLM3-3B-Base 上建立了一個具備神諭路由的受控 M-OPD 基準,將能力整合與路由模糊性分離開來。我們的研究揭示了顯著的能力整合差距:標準 M-OPD 相對於領域路由的神諭集成,僅能捕捉 35.6% 的可用改善空間,其中諸如指令遵循等簡潔任務更出現嚴重的性能衰退與過早停滯。關鍵的是,我們證明此一失敗並非源於梯度衝突,而是源於詞元級優化預算的嚴重錯配。此病理現象由三個正交因素驅動:跨領域的結構性序列長度差異、非均勻學習率導致的動態收斂漂移,以及非同步策略更新造成的多步獎勵陳舊。為解決這些失衡問題,我們提出了 Open-MOPD——一個整合詞元份額平衡、差距感知的動態預算分配以及學生獎勵刷新的原則性框架。這些機制共同系統性地恢復了跨領域平衡,將單一可部署學生的改善空間恢復率從 35.6% 提升至 83.4%。我們在學術界可負擔的硬體預算下,完全開源了端到端的後訓練方案、訓練軌跡與評估套件。
English
Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learning (RL) experts into a single generalist student via dense, token-level reward supervision. Despite its practical success, the optimization dynamics governing multi-teacher capability integration remain poorly understood, and open, rigorously reproducible recipes are conspicuously lacking. In this work, we establish a controlled M-OPD benchmark on SmolLM3-3B-Base with oracle routing, isolating capability integration from routing ambiguity. Our investigation reveals a pronounced capability integration gap: standard M-OPD captures only 35.6% of the available headroom relative to a domain-routed oracle ensemble, with concise tasks such as instruction following suffering severe degradation and premature stagnation. Crucially, we show that this failure stems not from gradient conflict, but from a severe misallocation of the token-level optimization budget. This pathology is driven by three orthogonal factors: structural sequence-length disparities across domains, dynamic convergence drift due to non-uniform learning rates, and multi-step reward staleness from asynchronous policy updates. To resolve these imbalances, we introduce Open-MOPD, a principled framework incorporating token-share balancing, gap-aware dynamic budget allocation, and student reward refresh. Together, these mechanisms systematically restore cross-domain balance, elevating headroom recovery from 35.6% to 83.4% in a single deployable student. We fully open-source our end-to-end post-training recipe, training trajectories, and evaluation suites on an academically accessible hardware budget.