影響導向蒸餾:解決取樣Token同策略蒸餾中的多樣性瓶頸
Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation
August 30, 2026
作者: Run Yang, Runpeng Dai, Jie Sun, Jielei Zhang, Fan Zhou, Hongtu Zhu, Peiyi Li, Longwen Gao
cs.AI
摘要
採樣詞元在策略蒸餾(OPD)能利用學生模型生成的詞元,有效地將能力從教師模型轉移至學生模型,且僅需教師模型對採樣詞元的概率。然而,它經常遭遇多樣性蒸餾失敗:學生的 pass@1 提升,但 pass@k 停滯不前,未能繼承教師模型的多樣性。為解釋此現象,我們引入了「一階局部熵影響力」(First-Order Local Entropy Influence),這是一個帶符號的一階代理指標,可將每次更新的熵效應分解為教師—學生對數概率差距與學生局部概率結構兩部分,並在經驗上將熵收縮與負影響位置連結起來。基於此,我們提出了「影響力引導的自適應在策略蒸餾」(IDA-OPD):它不依賴代價高昂的全詞彙前向 KL 散度目標,而是保留熵擴張的更新,並以散度自適應的優勢縮減取代熵收縮的更新,僅需教師模型的採樣詞元對數概率。在以推理為導向的蒸餾實驗中,IDA-OPD 持續提升 pass@k,透過蒸餾繼承教師模型的多樣性;以嚴格更低的成本達到最強教師資訊方法的表現,同時大致維持 vanilla OPD 的 pass@1,全程無需全詞彙教師資訊。
English
Sampled-token on-policy distillation (OPD) efficiently transfers capabilities from teacher to student using student-generated tokens, requiring teacher probabilities only for sampled tokens. Yet it frequently suffers from diversity distillation failure: the student's pass@1 improves while its pass@k plateaus, failing to inherit the teacher's diversity. To explain this, we introduce First-Order Local Entropy Influence, a signed first-order proxy that decouples each update's entropy effect into the teacher--student log-probability gap and the student's local probability structure, and empirically links entropy contraction to negative-influence positions. Motivated by this, we propose Influence-Directed Adaptive On-Policy Distillation (IDA-OPD): rather than relying on costly full-vocabulary Forward-KL objectives, it preserves entropy-expanding updates while replacing entropy-contracting ones with divergence-adaptive advantage shrinkage, using only the teacher's sampled-token log-probability. Experiments on reasoning-oriented distillation show IDA-OPD consistently improves pass@k, inheriting the teacher's diversity through distillation, matches the strongest teacher-informed methods at strictly lower cost, and broadly maintains vanilla OPD's pass@1, all without full-vocabulary teacher information.