Rebelse Student: Het Omkeren van Docentensignalen voor Redeneringsverkenning met Zelfgedistilleerde RLVR
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
May 11, 2026
Auteurs: Jeonghye Kim, Jiwon Jeon, Dongsheng Li, Yuqing Yang
cs.AI
Samenvatting
Zelf-distillatie is naar voren gekomen als een krachtig raamwerk voor het nabewerken van LLM's, waarbij een leraar, geconditioneerd op extra informatie, een student begeleidt zonder deze informatie, beide afkomstig van hetzelfde model. Hoewel deze begeleiding nuttig is wanneer de student faalt, overschrijft hetzelfde mechanisme bij succesvolle uitrolresultaten in plaats daarvan de keuzes van de student en onderdrukt het zijn eigen redenering. Daarom stellen we voor om het oorspronkelijke zelf-distillatiesignaal omgekeerd te lezen: wanneer de student slaagt op een pad dat de leraar niet zou hebben voorspeld, weerspiegelen deze tokens zijn zelfgestuurde redenering. Voortbouwend hierop stellen we RLRT (RLVR met Omgekeerde Leraar) voor, dat GRPO uitbreidt door deze tokens te versterken bij correcte uitrolresultaten. We interpreteren dit als een nieuwe vorm van verkenning in RLVR: geen uniforme diversiteit, maar waardevolle verkenning die geworteld is in het eigen succes van de student. Over basis-, instructie-getunede en denk-getunede Qwen3-checkpoints heen, presteert RLRT aanzienlijk beter dan zelf-distillatie en op verkenning gebaseerde basislijnen, waarmee informatie-asymmetrie wordt gevestigd als een nieuwe, principiële ontwerpas voor RLVR.
English
Self-distillation has emerged as a powerful framework for post-training LLMs, where a teacher conditioned on extra information guides a student without it, both from the same model. While this guidance is useful when the student has failed, on successful rollouts, the same mechanism instead overwrites the student's choices and suppresses it's own reasoning. Therefore, we propose reading the original self-distillation signal in reverse: when the student succeeds along a path the teacher would not have predicted, these tokens reflect its self-driven reasoning. Building on this, we propose RLRT (RLVR with Reversed Teacher), which augments GRPO by reinforcing these tokens on correct rollouts. We interpret this as a new form of exploration in RLVR: not uniform diversity, but valuable exploration grounded in the student's own success. Across base, instruction-tuned, and thinking-tuned Qwen3 checkpoints, RLRT substantially outperforms self-distillation and exploration-based baselines, establishing information asymmetry as a new, principled design axis for RLVR.