ChatPaper.aiChatPaper

Demystificatie van On-Policy Distillatie: Rollen, Pathologieën en Reguleringen

Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

July 15, 2026
Auteurs: Rui Wang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Wenhao Yu, Kam-Fai Wong
cs.AI

Samenvatting

On-policy distillatie (OPD) is uitgegroeid tot een sleutelparadigma in de post-training van LLM's, maar de trainingsdynamiek ervan blijft slecht begrepen. We presenteren een systematische studie naar de rol, pathologieën en regulering van OPD. Eerst verduidelijken we de rol van OPD als exploratiekatalysator: het stuurt de student via dichte token-niveau begeleiding naar correcte redeneerpaden, zonder de capaciteitsgrens te verleggen. We bevestigen dit door aan te tonen dat promptdiversiteit belangrijker is dan het aantal samples per probleem, en dat de effectiviteit van OPD volledig afhangt van de kwaliteit van het sturende signaal. Deze afhankelijkheid brengt twee pathologieën aan het licht die exploratie doen ontsporen. De Student-Docent Mismatch treedt op wanneer een grote verdelingkloof tussen docent en student ervoor zorgt dat het sturende signaal niet meer in lijn is met taakcorrectheid, waardoor exploratie in contraproductieve richtingen wordt gestuurd. Lengte-exploitatie ontstaat wanneer de geaggregeerde doelstelling op token-niveau lengteafhankelijke shortcuts creëert, waardoor de student het beloningslandschap kan bespelen via responsafkapping of overbodige padding, en zo ontaarde lengtemodi exploreert in plaats van redeneerstrategieën. Om deze pathologieën te beteugelen, onderzoeken we lichte signaalreguleringen: advantage clipping en log-schaalcompressie, die ervoor zorgen dat exploratie wordt geleid door betrouwbare signalen. Experimenten op zeven benchmarks tonen aan dat deze reguleringen lengte-exploitatie verminderen en effectieve distillatie mogelijk maken, waarbij ze stabiel beter presteren dan OPD-varianten en RLVR-baselines. Dit bevestigt dat niet louter de docentschaal, maar een goed gereguleerde signaalkwaliteit bepalend is voor succesvolle exploratie in OPD.
English
On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it steers the student toward correct reasoning paths via dense token-level guidance, without expanding capability ceiling. We confirm this by showing that prompt diversity matters more than per-problem sampling numbers, and critically, that the effectiveness of OPD hinges entirely on the quality of its guiding signal. This dependency exposes two pathologies that derail exploration. The Student-Teacher Mismatch occurs when a large teacher-student distributional gap causes the guiding signal to misalign with task correctness, steering exploration in counterproductive directions. Length Exploitation arises when the aggregated token-level objective creates length-dependent shortcuts, allowing the student to game the reward landscape through response truncation or redundant padding, exploring degenerate length modes rather than reasoning strategies. To tame these pathologies, we investigate lightweight signal regulations: advantage clipping and log-scale compression, ensuring exploration is guided by faithful signals. Experiments across seven benchmarks demonstrate that these regulations alleviate length exploitation and enable effective distillation, stably surpassing OPD variants and RLVR baselines, thereby confirming that well-regulated signal quality, rather than mere teacher scale, governs successful exploration in OPD.