직접 온-정책 증류를 통한 약-강 일반화

Weak-to-Strong Generalization via Direct On-Policy Distillation

July 8, 2026
저자: Shiyuan Feng, Huan-ang Gao, Haohan Chi, Hanlin Wu, Zhilong Zhang, Zheng Jiang, Bingxiang He, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
cs.AI

초록

검증 가능한 보상 기반 강화 학습(RLVR)은 언어 모델 추론을 향상시키는 강력한 방법이지만, 새로운 강력한 모델마다 반복 적용하기에는 비용이 많이 든다. 훈련 중 목표 모델이 많은 롤아웃을 생성해야 하기 때문이다. 모델 규모가 커짐에 따라 사후 훈련 자체가 병목 현상이 된다. 본 연구에서는 약-강 전이(weak-to-strong) 대안을 탐구한다: 롤아웃 비용이 더 저렴한 소형 모델에서 강화 학습을 수행한 후, 해당 강화 학습에서 얻은 지식을 재사용하여 더 강력한 목표 모델을 개선하는 것이다. 강화 학습 후의 약한 교사 모델을 직접 증류하는 것만으로는 충분하지 않다. 교사 모델의 최종 정책은 유용한 강화 학습 이득과 소형 모델의 한계가 혼합되어 있기 때문이다. 본 논문에서는 직접 정책 기반 증류(Direct On-Policy Distillation, Direct-OPD)를 제안한다. 이 방법은 교사 모델의 강화 학습 유도 정책 변화를 전이한다. Direct-OPD는 강화 학습 후의 교사 모델을 강화 학습 전의 동일 모델 기준과 비교하여, 그 로그 비율을 학생 모델을 위한 밀집 내재적 보상으로 활용한다. 쉽게 말해, 모델 쌍은 강화 학습이 약한 모델로 하여금 어떤 행동을 더 또는 덜 취하게 했는지를 알려주며, Direct-OPD는 이 신호를 강한 학생 모델의 자체 정책 기반 상태에 적용한다. 이는 약한 모델의 강화 학습 감독 신호를 직접 재사용하여, 목표 모델에 희소 보상 강화 학습을 수행할 필요를 없앤다. 실험적으로, Direct-OPD는 약한 교사 모델을 활용하여 강한 목표 모델을 일관되게 개선한다. 특히, AIME 2024에서 Qwen3-1.7B의 성능을 48.3%에서 58.3%로 향상시켰으며, 8대의 A100 GPU에서 단 4시간 만에 달성했다. 이는 단계별 직접 강화 학습보다 우수하며, 여러 정책 변화의 순차적 구성을 가능하게 한다. 본 결과는 강화 학습의 결과를 모방할 최종 모델이 아닌 내재적 보상 신호로서 모델 규모 간에 재사용할 수 있음을 보여준다.
English
Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to repeat on every new strong model because the target model must generate many rollouts during training. As models scale, post-training itself becomes a bottleneck. We study a weak-to-strong alternative: run RL on a smaller model where rollouts are cheaper, then reuse what that RL run learned to improve a stronger target model. Directly distilling the post-RL weak teacher is not enough, because the teacher's final policy mixes useful RL gains with the limitations of the smaller model. We propose Direct On-Policy Distillation (Direct-OPD), which transfers the teacher's RL-induced policy shift instead. Direct-OPD compares the post-RL teacher with its own pre-RL reference and treats their log-ratio as a dense implicit reward for the student. In plain terms, the checkpoint pair tells us which actions RL made the weak model more or less likely to take, and Direct-OPD applies that signal on the stronger student's own on-policy states. This directly reuses the weak model's RL supervision signal without running sparse-reward RL on the target model. Empirically, Direct-OPD consistently leverages weaker teachers to improve stronger target models; notably, it boosts Qwen3-1.7B from 48.3% to 58.3% on AIME 2024 in just 4 hours on 8 A100 GPUs. It outperforms step-matched direct RL and enables the sequential composition of multiple policy shifts. Our results show that RL outcomes can be reused across model scales as implicit reward signals, not merely as final models to imitate.
PDF932July 15, 2026