DOPD: Dubbele On-policy Distillatie
DOPD: Dual On-policy Distillation
June 29, 2026
Auteurs: Xinlei Yu, Gen Li, Qingyi Si, Guibin Zhang, Yuqi Xu, Congcong Wang, Shuai Dong, Kaiwen Tuo, Xiangyu Zeng, Kaituo Feng, Qunzhong Wang, Yang Shi, Xiaobin Hu, Xiangyu Yue, Jiaqi Wang, Shuicheng Yan
cs.AI
Samenvatting
On-policy distillatie (OPD) biedt superieure capaciteitsoverdracht door het superviseren van door de student gesamplede trajecten met dichte signalen op token-niveau. Om hoogwaardige supervisiebronnen te leveren en daarmee de prestatiefrontier van distillatie te verhogen, is een intuïtieve richting om bevoorrechte informatie aan de leraar of de student zelf toe te voegen. Echter, deze extra invoer induceert een potentiële faalmodus die we privilege-illusie noemen: een patroon dat de overdraagbare capaciteitskloof die studenten moeten dichten, verwart met de informatie-asymmetriekloof die alleen kan worden nagebootst maar nooit gerepliceerd. Dit probleem wordt verder versterkt door de inherente niet-uniformiteit van token-niveau supervisie, waarbij slechts een kleine subset van tokens cruciale capaciteitsdragende signalen draagt. Daartoe stellen we DOPD voor, een voordeelbewust duaal distillatieparadigma dat dynamisch token-niveau supervisie routeert tussen bevoorrechte leraar- en bevoorrechte student-beleid, gebaseerd op hun voordeelkloof en relatieve kansen. Elk token ontvangt supervisie van verschillende sterkte, doelstelling en strategie van ofwel de leraar ofwel de student zelf, wat geloofwaardige capaciteit overdraagt terwijl het tegelijkertijd hulpsignalen ontvangt, om privilege-illusie te verlichten. Uitgebreide experimenten op zowel grote taalmodellen (LLM) als visie-taalmodellen (VLM) tonen aan dat DOPD consistent beter presteert dan Vanilla OPD en andere tegenhangers. Verdere resultaten op stabiliteit, robuustheid, continu leren en out-of-distribution taken valideren zijn superioriteit.
English
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sources and thereby elevate the performance frontier of distillation, an intuitive direction is to infuse privileged information to either teacher or student itself. However, this additional input induces a potential failure mode we dub privilege illusion: a pattern that conflates the transferable capability gap that students are meant to close, and the information asymmetry gap that can only be mimicked but never replicated. This issue is further amplified by the inherent non-uniformity of token-level supervision, where only a small subset of tokens carries pivotal capability-bearing signals. To this end, we propose DOPD, an advantage-aware dual distillation paradigm that dynamically routes token-level supervision between privileged teacher and privileged student policies based on their advantage gap and relative probabilities. Each token receives supervision of different strength, objective, and strategy from either teacher or student itself, which transfers credible capability while simultaneously receiving auxiliary signals, to alleviate privilege illusion. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that DOPD consistently outperforms Vanilla OPD and other counterparts. Further results on stability, robustness, continual learning, and out-of-distribution tasks validate its superiority.