ChatPaper.aiChatPaper

Open-MOPD:マルチティーチャー・オンポリシー蒸留における能力不均衡の診断と修正

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

August 19, 2026
著者: Huan-ang Gao, Haohan Chi, Yong Yan, Shiyuan Feng, Hanlin Wu, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
cs.AI

要旨

マルチ教師オン・ポリシー蒸留(M-OPD)は、密なトークンレベルの報酬監視を通じて、ドメイン特化型の強化学習(RL)エキスパート群を単一の汎用学生モデルへ統合する有望なパラダイムとして登場した。その実用的な成功にもかかわらず、マルチ教師による能力統合を支配する最適化ダイナミクスは十分に理解されておらず、厳密に再現可能な公開レシピも顕著に欠如している。本研究では、オラクル・ルーティングを備えたSmolLM3-3B-Base上に制御されたM-OPDベンチマークを構築し、能力統合をルーティングの曖昧性から切り離して評価する。我々の調査は、顕著な能力統合ギャップを明らかにした。標準的なM-OPDは、ドメイン・ルーティング型オラクル・アンサンブルに対する利用可能なヘッドルーム(改善余地)のわずか35.6%しか獲得できず、指示追従などの簡潔なタスクでは深刻な性能低下と早期停滞が生じる。重要なことに、この失敗は勾配競合ではなく、トークンレベルの最適化予算の深刻な誤配分に起因することを示す。この病態は、3つの直交する要因によって引き起こされる。すなわち、ドメイン間の構造的な系列長の不均一性、非一様な学習率による動的収束ドリフト、そして非同期ポリシー更新による多段階報酬の陳腐化である。これらの不均衡を解消するため、我々はトークン共有バランシング、ギャップ認識型動的予算配分、学生報酬リフレッシュを統合した原理的なフレームワークであるOpen-MOPDを導入する。これらのメカニズムは連携して、クロスドメインのバランスを体系的に回復させ、単一のデプロイ可能な学生モデルにおいてヘッドルーム回復率を35.6%から83.4%へと引き上げる。我々は、学術的にアクセス可能なハードウェア予算の下で、エンドツーエンドのポストトレーニングレシピ、トレーニング軌跡、評価スイートを完全にオープンソース化する。
English
Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learning (RL) experts into a single generalist student via dense, token-level reward supervision. Despite its practical success, the optimization dynamics governing multi-teacher capability integration remain poorly understood, and open, rigorously reproducible recipes are conspicuously lacking. In this work, we establish a controlled M-OPD benchmark on SmolLM3-3B-Base with oracle routing, isolating capability integration from routing ambiguity. Our investigation reveals a pronounced capability integration gap: standard M-OPD captures only 35.6% of the available headroom relative to a domain-routed oracle ensemble, with concise tasks such as instruction following suffering severe degradation and premature stagnation. Crucially, we show that this failure stems not from gradient conflict, but from a severe misallocation of the token-level optimization budget. This pathology is driven by three orthogonal factors: structural sequence-length disparities across domains, dynamic convergence drift due to non-uniform learning rates, and multi-step reward staleness from asynchronous policy updates. To resolve these imbalances, we introduce Open-MOPD, a principled framework incorporating token-share balancing, gap-aware dynamic budget allocation, and student reward refresh. Together, these mechanisms systematically restore cross-domain balance, elevating headroom recovery from 35.6% to 83.4% in a single deployable student. We fully open-source our end-to-end post-training recipe, training trajectories, and evaluation suites on an academically accessible hardware budget.