ChatPaper.aiChatPaper

온-정책 증류 명료화: 역할, 병리, 및 규제

Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

July 15, 2026
저자: Rui Wang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Wenhao Yu, Kam-Fai Wong
cs.AI

초록

온정책 증류(On-policy Distillation, OPD)는 LLM 후훈련의 핵심 패러다임으로 자리 잡았지만, 그 훈련 동역학은 여전히 제대로 이해되지 않고 있다. 본 연구에서는 OPD의 역할, 병리 현상 및 조정 메커니즘을 체계적으로 조사한다. 먼저 OPD의 역할을 탐색 촉매로 명확히 규명한다. OPD는 밀집된 토큰 수준의 지침을 통해 학생 모델을 올바른 추론 경로로 유도하되, 능력 상한선을 확장하지는 않는다. 이는 문제별 샘플링 수보다 프롬프트 다양성이 더 중요하며, 결정적으로 OPD의 효과가 안내 신호의 질에 전적으로 의존한다는 점을 입증함으로써 확인된다. 이러한 의존성은 탐색을 방해하는 두 가지 병리 현상을 드러낸다. 학생-교사 불일치(Student-Teacher Mismatch)는 큰 교사-학생 분포 차이로 인해 안내 신호가 과제 정확성과 불일치하게 되어 역효과를 내는 방향으로 탐색을 유도할 때 발생한다. 길이 악용(Length Exploitation)은 집계된 토큰 수준 목표 함수가 길이에 의존적인 지름길을 만들어 학생이 응답 절단이나 불필요한 패딩을 통해 보상 환경을 조작하게 하여, 추론 전략 대신 퇴행적인 길이 모드를 탐색하게 할 때 발생한다. 이러한 병리 현상을 제어하기 위해 경량의 신호 조정 방법, 즉 어드밴티지 클리핑과 로그 스케일 압축을 조사하여 탐색이 신뢰할 수 있는 신호에 의해 안내되도록 한다. 일곱 개의 벤치마크에 걸친 실험은 이러한 조정이 길이 악용을 완화하고 효과적인 증류를 가능하게 하여 OPD 변형 및 RLVR 기준선을 안정적으로 능가함을 보여준다. 이는 OPD에서 성공적인 탐색을 좌우하는 요소가 단순한 교사 규모가 아니라 잘 조정된 신호 품질임을 확인한다.
English
On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it steers the student toward correct reasoning paths via dense token-level guidance, without expanding capability ceiling. We confirm this by showing that prompt diversity matters more than per-problem sampling numbers, and critically, that the effectiveness of OPD hinges entirely on the quality of its guiding signal. This dependency exposes two pathologies that derail exploration. The Student-Teacher Mismatch occurs when a large teacher-student distributional gap causes the guiding signal to misalign with task correctness, steering exploration in counterproductive directions. Length Exploitation arises when the aggregated token-level objective creates length-dependent shortcuts, allowing the student to game the reward landscape through response truncation or redundant padding, exploring degenerate length modes rather than reasoning strategies. To tame these pathologies, we investigate lightweight signal regulations: advantage clipping and log-scale compression, ensuring exploration is guided by faithful signals. Experiments across seven benchmarks demonstrate that these regulations alleviate length exploitation and enable effective distillation, stably surpassing OPD variants and RLVR baselines, thereby confirming that well-regulated signal quality, rather than mere teacher scale, governs successful exploration in OPD.