當教師誤導:偽訊號感知的同策略蒸餾
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
August 4, 2026
作者: Yinuo Jiang, Yongjie Ye, Zhou Tao, Xiang Zhuang, Qiang Zhang, Huajun Chen, Tiankai Li
cs.AI
摘要
同策略蒸餾(On-Policy Distillation, OPD)透過以密集的詞元層級教師訊號監督學生取樣的軌跡,來轉移教師能力。近年來的選擇性同策略蒸餾方法透過優先選擇具有信心、資訊量或可學習性的訊號來改善此流程。然而,這些假設忽略了語言模型一個根本的失效模式:其詞元層級的判斷可能受到與輸入無關的語言先驗、格式慣例或刻板推理模板所驅動,而非任務特定的證據。我們將此類與最佳化相關但輸入扎根程度薄弱的監督訊號稱為同策略蒸餾中的虛假訊號(spurious signals),它們可能產生較大的梯度,卻對任務改進方向貢獻甚微。為緩解此問題,我們提出 SA-OPD,一個基於輸入扎根程度與最佳化影響來識別並過濾誤導性詞元層級監督的虛假訊號感知同策略蒸餾框架(Spurious-Signal-Aware On-Policy Distillation)。SA-OPD 引入了一個輕量級的輸入扎根程度代理指標,用以估計詞元層級的蒸餾訊號是否真正依賴於輸入。接著,它僅過濾掉同時表現出低輸入扎根程度與極端蒸餾分歧的詞元,從而移除高影響力的虛假更新,達成細粒度的同策略蒸餾最佳化。在大型語言模型(LLM)與視覺語言模型(VLM)設定上的大量實驗顯示,SA-OPD 持續優於基礎同策略蒸餾(Vanilla OPD)及具競爭力的選擇性方法。這些結果確立了輸入扎根程度作為同策略蒸餾監督選擇的關鍵維度,並提供了一種簡單而有效的策略來緩解虛假更新。
English
On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable. However, the assumptions overlook a fundamental failure mode of language models: their token-level judgments can be driven by input-agnostic language priors, formatting conventions, or stereotyped reasoning templates rather than task-specific evidence. We refer to such optimization-relevant but weakly input-grounded supervision as spurious signals in OPD, which may produce large gradients while contributing little task-improving direction. To mitigate this issue, we propose SA-OPD, a Spurious-Signal-Aware On-Policy Distillation framework that identifies and filters misleading token-level supervision based on input-groundedness and optimization impact. SA-OPD introduces a lightweight input-groundedness proxy estimating whether a token-level distillation signal truly depends on the input. It then filters only tokens that simultaneously exhibit low input-groundedness and extreme distillation divergence, thereby removing high-impact spurious updates and achieving fine-grained OPD optimization. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that SA-OPD consistently outperforms Vanilla OPD and competitive selective methods. These results establish input-groundedness as a key dimension for OPD supervision selection and offer a simple, effective strategy for mitigating spurious updates.