信賴區域策略蒸餾
Trust Region Policy Distillation
July 6, 2026
作者: Zhengpeng Xie, Li Lyna Zhang, Zeke Xie, Mao Yang
cs.AI
摘要
宏大目標難以一蹴而就,將其分解為小步驟更為明智。我們提出了信任區域策略蒸餾(Trust Region Policy Distillation, TOP-D),通過動態構建近端教師,將原本不穩定、高方差的同策略蒸餾(On-Policy Distillation, OPD)轉化為穩定的訓練範式。理論上,我們建立了一個嚴謹的框架,證明TOP-D能從本質上控制梯度方差。通過提供正式的全局收斂分析和單調改進界限,我們從數學上形式化了整體訓練動態的可靠性與穩定性。實證方面,TOP-D在數學推理任務中顯著提升了訓練穩定性、樣本效率以及最終性能。更重要的是,TOP-D不引入任何額外計算開銷,使其成為成熟OPD範式的一個極具前景的替代方案。
English
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.