검증 가능한 보상을 사용한 온-정책 증류
On-policy Distillation with Verifiable Reward
August 25, 2026
저자: Wenze Lin, Jiale Zhao, Xitai Jiang, Songde Rao, Yining Li, Shenzhi Wang, Bingxiang He, Gao Huang
cs.AI
초록
검증 가능한 보상을 사용한 강화 학습(RLVR)과 온-정책 증류(OPD)는 대규모 언어 모델의 사후 훈련을 위해 널리 채택된 두 가지 패러다임이 되었다. 그러나 RLVR은 희소한 과제 수준 피드백으로 인해 어려움을 겪는 반면, OPD는 조밀한 토큰 수준 안내를 제공하지만 궤적 정확성을 무시하여 성능이 교사 모델 수준으로 제한된다. 이 둘을 결합하는 것은 유망한 방향이다. OPD는 조밀한 감독 신호를 제공하고 RLVR은 과제 수준의 정확성을 제공하기 때문이다. 그럼에도 불구하고 기존의 통합 방식은 종종 가중 결합이나 휴리스틱 전환에 의존하여 추가 하이퍼파라미터와 절충을 도입한다. 우리는 추가 하이퍼파라미터 없이 OPD와 RLVR을 매끄럽게 결합하는 간단하면서도 효과적인 방법인 검증 가능한 보상을 사용한 온-정책 증류(OPDVR)를 제안한다. 우리는 먼저 궤적 정확성에 기반하여 샘플링된 토큰 OPD의 암묵적 보상을 재구성한 다음, ReLU 게이팅 메커니즘을 적용하여 올바른 궤적은 음이 아닌 보상을 받고 잘못된 궤적은 양이 아닌 보상을 받도록 한다. 이를 통해 교사의 분포적 안내를 유지하면서 증류 신호를 과제 성공에 정렬한다. 더 나아가, 우리의 수정은 샘플링된 토큰 OPD를 적절한 RLVR 방법으로 변환하여 GRPO와 같은 모든 정책 경사 알고리즘과 쉽게 결합할 수 있게 한다. 여섯 가지 추론 벤치마크에 대한 실험은 OPDVR이 표준 OPD보다 일관되게 우수한 성능을 보임을 보여준다. 우리의 코드는 https://github.com/LeapLabTHU/OPDVR에서 확인할 수 있다.
English
Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled-token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards---thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled-token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.