모방을 넘어서: 추론 진행에 따른 온-폴리시 증류 필터링
Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress
August 19, 2026
저자: Chen Yang, Haiyuan Wan, Rengrong Xiong, Yize Chen, Danny H. K. Tsang
cs.AI
초록
온-정책 증류(OPD)는 학생이 생성한 궤적을 교사의 밀집된 토큰 수준 감독과 짝지음으로써 언어 모델의 사후 훈련을 위한 효과적인 프레임워크로 부상했다. 그러나 OPD는 교사로부터 도출된 보상이 추론 진행의 적절한 대리 지표라고 암묵적으로 가정하며, 따라서 정책 최적화 중 모든 교사 피드백을 동등하게 취급한다. 그러나 실제로 이 가정은 항상 성립하지 않는다. 우리는 교사 유래 보상이 실제 추론 진행과 자주 충돌함을 관찰한다. 명확한 추론 발전이 있는 추론 단계라도 단순히 교사의 출력과의 편차 때문에 더 낮은 증류 보상을 받을 수 있다. 이러한 불일치를 해결하기 위해, 우리는 온-정책 증류를 위한 추론 진행 인식 보상 필터링(R2-OPD)을 제안한다. 이 방법은 추론 스팬에 대해 궤적 내 두 가지 순위를 구성하는데, 하나는 교사 유래 보상에서, 다른 하나는 독립적으로 추정된 진행 보상에서 얻는다. 두 순위가 일치하지 않을 때마다 증류 보상을 선택적으로 억제하여, 효과적인 교사 지도를 유지하면서 추론 진행과 충돌하는 감독을 줄인다. 우리의 접근 방식은 특히 추론 성능 측면에서 표준 OPD보다 일관된 개선을 보여준다.
English
On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing student-generated trajectories with dense token-level supervision from a teacher. However, OPD implicitly assumes that teacher-derived rewards are an appropriate proxy for reasoning progress, and therefore treats all teacher feedback equally during policy optimization. While in practice, this assumption does not always hold. We observe that teacher-derived rewards often conflict with genuine reasoning progress, as reasoning steps with clear reasoning advancement may still receive lower distillation rewards, simply due to deviation from teacher's outputs. To address this mismatch, we propose Reasoning-Progress-Aware Reward Filtering for On-Policy Distillation (R2-OPD), which constructs two within-trajectory rankings of reasoning spans, one from teacher-derived rewards and the other from independently estimated progress reward. Distillation rewards are selectively suppressed whenever the two rankings disagree, reducing supervision that conflicts with reasoning progress while preserving effective teacher guidance. Our approach shows consistent improvement over standard OPD especially regarding reasoning performances.