ChatPaper.aiChatPaper

信頼領域ポリシー蒸留

Trust Region Policy Distillation

July 6, 2026
著者: Zhengpeng Xie, Li Lyna Zhang, Zeke Xie, Mao Yang
cs.AI

要旨

大きな目標を一度に達成するのは難しく、小さなステップに分割する方が賢明です。本稿では、信頼領域方策蒸留(TOP-D)を提案します。これは、動的に近接教師を構築することで、不安定性と高分散で有名なオンポリシー蒸留(OPD)を安定した学習パラダイムへと変換します。理論的には、TOP-Dが本質的に勾配分散を制御することを示す厳密な枠組みを確立します。形式的な大域的収束解析と単調改善境界を提供することで、学習ダイナミクス全体の信頼性と安定性を数学的に形式化します。実験的には、TOP-Dは数学的推論タスクにおいて、学習の安定性、サンプル効率、および最終性能を劇的に向上させます。さらに重要なことに、TOP-Dは追加の計算オーバーヘッドを一切導入せず、確立されたOPDパラダイムに対する有望な代替手法として位置づけられます。
English
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.