ChatPaper.aiChatPaper

신뢰 영역 정책 증류

Trust Region Policy Distillation

July 6, 2026
저자: Zhengpeng Xie, Li Lyna Zhang, Zeke Xie, Mao Yang
cs.AI

초록

큰 목표를 한 번에 달성하기는 어렵고, 이를 작은 단계로 나누는 것이 더 현명하다. 본 논문에서는 동적으로 근접 교사를 구성함으로써 악명 높은 불안정성과 고분산 특성을 지닌 온-정책 증류(OPD)를 안정적인 훈련 패러다임으로 변환하는 신뢰 영역 정책 증류(TOP-D)를 제안한다. 이론적으로, 우리는 TOP-D가 본질적으로 그래디언트 분산을 제어함을 보여주는 엄밀한 프레임워크를 확립한다. 단조 개선 한계와 함께 형식적 전역 수렴 분석을 제공함으로써, 전체 훈련 동역학의 신뢰성과 안정성을 수학적으로 공식화한다. 실증적으로, TOP-D는 수학적 추론 과제에서 훈련 안정성, 샘플 효율성 및 최종 성능을 극적으로 향상시킨다. 더 중요한 것은, TOP-D가 추가 계산 오버헤드를 전혀 도입하지 않는다는 점이며, 이는 잘 확립된 OPD 패러다임에 대한 유망한 대안으로 자리매김한다.
English
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.