ChatPaper.aiChatPaper

以同策略反向蒸餾誘發弱至強泛化

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

September 8, 2026
作者: Youngrok Park, Sangmin Bae, Hojung Jung, Jongwoo Ko, Yunseon Choi, Young Jin Kim, Pashmina Cameron, Aaron Courville, Se-Young Yun
cs.AI

摘要

弱到強泛化探討較強模型能否向較弱監督者學習並超越它們。這個問題對連續模型世代與多領域整合尤其重要,因為從頭重複進行前沿規模的後訓練可能成本高昂到難以承受。然而,傳統蒸餾將弱教師視為最佳化目標,可能將其容量上限加諸於學生。我們提出同策略反向蒸餾(On-Policy Reverse Distillation, OPRD),其在學生 rollout 上評估教師相對於其參考策略的策略偏移,並沿該方向放大學生由驗證器驅動的策略梯度分量。OPRD 僅重新縮放驗證器支援的更新,因而保留策略最佳化的駐點,同時加速學習以超越教師。在連續模型轉移與多教師蒸餾中,OPRD 均以比現有強化學習(RL)與蒸餾方法更少的學生更新次數,達到更高效能。回應風格分析顯示,相較於其弱教師,OPRD 學生仍更接近僅使用基於驗證器的強化學習訓練的模型,意味著教師引導加速而非重新導向學生自身的最佳化。傳統強到弱蒸餾的結果進一步證明,無論容量排序為何,OPRD 都能有效結合驗證器驅動的策略最佳化與教師引導。
English
Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet conventional distillation treats the weak teacher as an optimization target, potentially imposing its capacity ceiling on the student. We introduce On-Policy Reverse Distillation (OPRD), which evaluates the teacher's policy shift relative to its reference policy on student rollouts and amplifies the component of the student's verifier-driven policy gradient along that direction. By rescaling only verifier-supported updates, OPRD preserves the stationary points of policy optimization while accelerating learning beyond the teacher. In both successive model transfer and multi-teacher distillation, OPRD achieves higher performance with fewer student updates than existing RL and distillation approaches. Response-style analysis shows that OPRD students remain closer to models trained with verifier-based RL alone than to their weak teachers, suggesting that teacher guidance accelerates rather than redirects the student's own optimization. Results in conventional strong-to-weak distillation further demonstrate that OPRD effectively combines verifier-driven policy optimization with teacher guidance regardless of capacity ordering.