ChatPaper.aiChatPaper

Trust Region Beleidsdestillatie

Trust Region Policy Distillation

July 6, 2026
Auteurs: Zhengpeng Xie, Li Lyna Zhang, Zeke Xie, Mao Yang
cs.AI

Samenvatting

Grote doelen zijn moeilijk in één keer te bereiken; het is verstandiger om ze op te splitsen in kleine stappen. Wij presenteren Trust Region Policy Distillation (TOP-D), dat de beruchte instabiele en hoog-variante On-Policy Distillation (OPD) omvormt tot een stabiel trainingsparadigma door dynamisch een proximale leraar te construeren. Theoretisch leggen we een rigoureus kader vast dat aantoont dat TOP-D inherent de gradiëntvariantie beheerst. Door een formele globale convergentieanalyse samen met een monotone verbeteringsgrens te bieden, formaliseren we wiskundig de betrouwbaarheid en stabiliteit van de algehele trainingsdynamiek. Empirisch gezien verbetert TOP-D de trainingsstabiliteit, sample-efficiëntie en uiteindelijke prestaties bij wiskundige redeneertaken aanzienlijk. Nog belangrijker is dat TOP-D nul extra rekenkundige overhead introduceert, wat het positioneert als een veelbelovend alternatief voor het gevestigde OPD-paradigma.
English
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.