온폴리시 역증류를 이용한 약-강 일반화 유도
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
September 8, 2026
저자: Youngrok Park, Sangmin Bae, Hojung Jung, Jongwoo Ko, Yunseon Choi, Young Jin Kim, Pashmina Cameron, Aaron Courville, Se-Young Yun
cs.AI
초록
약-강 일반화는 더 강한 모델이 더 약한 감독자로부터 학습하여 그들을 능가할 수 있는지 묻는다. 이 질문은 프런티어 규모의 사후 학습을 처음부터 반복하는 것이 감당하기 어려울 정도로 비쌀 수 있는 연속적 모델 세대와 다중 도메인 통합에서 특히 중요하다. 그러나 기존 증류는 약한 교사를 최적화 목표로 취급하여, 잠재적으로 학생에게 교사의 용량 상한을 부과한다. 우리는 온-폴리시 역증류(On-Policy Reverse Distillation, OPRD)를 소개한다. 이는 학생 롤아웃에서 교사의 정책이 참조 정책에 대해 얼마나 이동했는지를 평가하고, 학생의 검증기 기반 정책 그래디언트 중 그 방향에 해당하는 성분을 증폭한다. 검증기가 뒷받침하는 업데이트만 재스케일링함으로써, OPRD는 정책 최적화의 정류점을 보존하면서 교사를 넘어서는 학습을 가속화한다. 연속적 모델 전이와 다중 교사 증류 모두에서 OPRD는 기존 RL 및 증류 접근법보다 더 적은 학생 업데이트로 더 높은 성능을 달성한다. 응답 스타일 분석은 OPRD 학생들이 약한 교사보다 검증기 기반 RL만으로 학습된 모델에 더 가깝게 남아 있음을 보여주며, 이는 교사의 지도가 학생 자체의 최적화를 다른 방향으로 전환하기보다 가속화함을 시사한다. 기존 강-약 증류에서의 결과는 OPRD가 용량 순서와 관계없이 검증기 기반 정책 최적화와 교사 지도를 효과적으로 결합함을 추가로 입증한다.
English
Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet conventional distillation treats the weak teacher as an optimization target, potentially imposing its capacity ceiling on the student. We introduce On-Policy Reverse Distillation (OPRD), which evaluates the teacher's policy shift relative to its reference policy on student rollouts and amplifies the component of the student's verifier-driven policy gradient along that direction. By rescaling only verifier-supported updates, OPRD preserves the stationary points of policy optimization while accelerating learning beyond the teacher. In both successive model transfer and multi-teacher distillation, OPRD achieves higher performance with fewer student updates than existing RL and distillation approaches. Response-style analysis shows that OPRD students remain closer to models trained with verifier-based RL alone than to their weak teachers, suggesting that teacher guidance accelerates rather than redirects the student's own optimization. Results in conventional strong-to-weak distillation further demonstrate that OPRD effectively combines verifier-driven policy optimization with teacher guidance regardless of capacity ordering.