ChatPaper.aiChatPaper

DiffGate:基于难度门控的在线蒸馏教师指导

DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation

October 3, 2026
作者: Karn Tiwari, Varnith Chordia, Prathosh A P
cs.AI

摘要

在线策略蒸馏(OPD)已成为大语言模型后训练中广泛使用的范式,它通过在学生模型自身生成的轨迹上对其进行监督,减少了传统蒸馏中的训练-测试不匹配。然而,现有的 OPD 目标在很大程度上仍是词元局部的且与结果无关,尽管推理质量是在轨迹层面决定的,却仍在每个前缀上优化教师-学生一致性。可验证奖励的强化学习(RLVR),尤其是组相对策略优化(GRPO),提供了互补的结果级监督,但存在奖励稀疏和信用分配粗糙的问题。我们表明,OPD 与 RLVR 展现出互补的盲点:教师信号提供密集的局部指导,但与轨迹正确性弱对齐;而组相对奖励能捕捉任务成功,却提供粗糙的词元级信用,并在全失败组中消失。我们提出 DiffGate,一种结果门控目标,将 GRPO 与选择性、有界的教师指导相结合。教师监督仅应用于失败轨迹,按组难度缩放,并进行平滑有界处理,以防止极端的教师-学生差异主导优化。因此,验证器决定哪些轨迹接受教师指导,而教师在这些轨迹内提供密集的词元级更新方向。在 Qwen3-0.6B 和 Qwen3-1.7B 学生模型上,DiffGate 相比匹配的 GRPO 将代码 avg@8 分别提升 +1.7 和 +1.8 分,将 pass@8 分别提升 +1.6 和 +5.7 分。在数学任务上,avg@8 与 GRPO 的差距保持在 0.5 分以内,而 pass@8 分别提升 +1.1 和 +3.9 分。总体而言,DiffGate 在所有四种模型-领域设置中均提升了 pass@8,表明在我们的评估协议下解覆盖范围得到提高。
English
On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely token-local and outcome-agnostic, optimizing teacher--student agreement at each prefix despite reasoning quality being determined at the trajectory level. Reinforcement learning with verifiable rewards (RLVR), particularly Group Relative Policy Optimization (GRPO), provides complementary outcome-level supervision but suffers from sparse rewards and coarse credit assignment. We show that OPD and RLVR exhibit complementary blind spots: teacher signals provide dense local guidance but are weakly aligned with rollout correctness, whereas group-relative rewards capture task success but provide coarse token-level credit and vanish on all-failure groups. We introduce DiffGate, an outcome-gated objective that combines GRPO with selective, bounded teacher guidance. Teacher supervision is applied only to failed trajectories, scaled by group difficulty, and smoothly bounded to prevent extreme teacher--student discrepancies from dominating optimization. The verifier therefore determines which trajectories receive teacher guidance, while the teacher provides dense token-level update directions within those trajectories. Across Qwen3-0.6B and Qwen3-1.7B students, DiffGate improves code avg@8 over matched GRPO by +1.7 and +1.8 points and pass@8 by +1.6 and +5.7 points, respectively. On mathematics, avg@8 remains within 0.5 points of GRPO while pass@8 improves by +1.1 and +3.9 points. Overall, DiffGate improves pass@8 across all four model--domain settings, demonstrating improved solution coverage under our evaluation protocol.