教師が誤解を招くとき:疑似信号を考慮したオンポリシー蒸留
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
August 4, 2026
著者: Yinuo Jiang, Yongjie Ye, Zhou Tao, Xiang Zhuang, Qiang Zhang, Huajun Chen, Tiankai Li
cs.AI
要旨
オン・ポリシー蒸留(On-Policy distillation, OPD)は、生徒モデルがサンプリングした軌跡を、高密度なトークン単位の教師信号で監視することにより、教師の能力を転移する。近年の選択的OPD手法は、確信度が高い、情報量が多い、または学習可能性が高い信号を優先することによって、このプロセスを改善している。しかしながら、これらの仮定は、言語モデルの根本的な失敗モードを見落としている。すなわち、言語モデルのトークン単位の判断は、タスク固有の証拠ではなく、入力に依存しない言語事前分布、書式規則、または定型化された推論テンプレートによって駆動される可能性がある。我々は、このような最適化には関連するものの入力接地性が弱い教師信号を、OPDにおけるスプリアス信号(疑似信号)と呼ぶ。これらの信号は、大きな勾配を生み出す一方で、タスク改善の方向への寄与はほとんどない。この問題を緩和するために、我々はSA-OPD(Spurious-Signal-Aware On-Policy Distillation)フレームワークを提案する。これは、入力接地性と最適化への影響に基づいて、誤解を招くトークン単位の教師信号を特定し、フィルタリングする。SA-OPDは、トークン単位の蒸留信号が本当に入力に依存しているかどうかを推定する軽量な入力接地性プロキシを導入する。そして、入力接地性が低く、かつ蒸留ダイバージェンスが極端に大きいトークンのみをフィルタリングする。これにより、影響の大きなスプリアス更新を除去し、細粒度のOPD最適化を実現する。大規模言語モデル(LLM)と視覚言語モデル(VLM)の両設定における広範な実験により、SA-OPDがバニラOPDおよび競争力のある選択的手法を一貫して上回ることを実証する。これらの結果は、入力接地性がOPDの教師信号選択における重要な次元であることを確立し、スプリアス更新を軽減するためのシンプルで効果的な戦略を提供する。
English
On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable. However, the assumptions overlook a fundamental failure mode of language models: their token-level judgments can be driven by input-agnostic language priors, formatting conventions, or stereotyped reasoning templates rather than task-specific evidence. We refer to such optimization-relevant but weakly input-grounded supervision as spurious signals in OPD, which may produce large gradients while contributing little task-improving direction. To mitigate this issue, we propose SA-OPD, a Spurious-Signal-Aware On-Policy Distillation framework that identifies and filters misleading token-level supervision based on input-groundedness and optimization impact. SA-OPD introduces a lightweight input-groundedness proxy estimating whether a token-level distillation signal truly depends on the input. It then filters only tokens that simultaneously exhibit low input-groundedness and extreme distillation divergence, thereby removing high-impact spurious updates and achieving fine-grained OPD optimization. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that SA-OPD consistently outperforms Vanilla OPD and competitive selective methods. These results establish input-groundedness as a key dimension for OPD supervision selection and offer a simple, effective strategy for mitigating spurious updates.