Destilação de Política por Região de Confiança
Trust Region Policy Distillation
July 6, 2026
Autores: Zhengpeng Xie, Li Lyna Zhang, Zeke Xie, Mao Yang
cs.AI
Resumo
Grandes objetivos são difíceis de alcançar de uma só vez; dividi-los em pequenos passos é mais sábio. Apresentamos a Destilação de Política por Região de Confiança (TOP-D), que transforma a notoriamente instável e de alta variância Destilação On-Policy (OPD) em um paradigma de treinamento estável, construindo dinamicamente um professor proximal. Teoricamente, estabelecemos um arcabouço rigoroso demonstrando que o TOP-D controla inerentemente a variância do gradiente. Ao fornecer uma análise formal de convergência global juntamente com um limite de melhoria monotônica, formalizamos matematicamente a confiabilidade e a estabilidade da dinâmica geral de treinamento. Empiricamente, o TOP-D melhora drasticamente a estabilidade do treinamento, a eficiência de amostragem e o desempenho final em tarefas de raciocínio matemático. Mais importante, o TOP-D introduz zero de carga computacional adicional, posicionando-se como uma alternativa promissora ao bem estabelecido paradigma OPD.
English
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.