信任区域策略蒸馏
Trust Region Policy Distillation
July 6, 2026
作者: Zhengpeng Xie, Li Lyna Zhang, Zeke Xie, Mao Yang
cs.AI
摘要
宏大目标难以一蹴而就,分步实现更为明智。我们提出信任区域策略蒸馏(Trust Region Policy Distillation, TOP-D),通过动态构建邻近教师(proximal teacher),将原本极不稳定、高方差的在线策略蒸馏(On-Policy Distillation, OPD)转化为稳定的训练范式。理论层面,我们构建了一个严谨的框架,证明TOP-D能够内在地控制梯度方差;通过给出形式化的全局收敛分析与单调提升界,我们从数学上形式化了整体训练动态的可靠性与稳定性。实验层面,TOP-D在数学推理任务上显著提升了训练稳定性、样本效率与最终性能。更重要的是,TOP-D不引入任何额外计算开销,这使其成为成熟OPD范式的一个有前景的替代方案。
English
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.