每枚硬幣皆有兩面:論大型語言模型在線策略蒸餾中泛化的雙重性
Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models
August 17, 2026
作者: Zhaoyi Li, Deyang Kong, Yuan Wei, Evan Yang, Ranran Shen, Mahardika Krisna Ihsani, Ming Yang, Wei Zhang, Chuan Hao, Jian Yang, Ran Tao, Bryan Dai, Shikun Zhang, Wei Ye, Ying Wei, Defu Lian
cs.AI
摘要
在策略蒸餾(On-policy distillation, OPD)透過監督來自學生自身策略所採樣的軌跡來遷移教師能力,然而其泛化行為至今仍未被充分理解,因為多數研究僅在單一領域及鄰近訓練資料的基準上評估OPD。我們提出了一個受控研究,每次僅變動單一泛化因素,範圍從域內分布偏移到跨域遷移及多教師設定。我們發現OPD遷移的是教師的推理行為,而非其對特定問題的答案:訓練難度幾乎不影響結果,甚至教師從未解出的問題也具有效用。遷移效果強烈取決於教師與學生之間的來源關係:同來源配對能使學生在跨語言、推理範圍、甚至其他領域上都接近教師,而跨來源配對則大多僅是擬合訓練分布。這種廣泛的影響力是一把雙刃劍:由於將提示路由至領域專家無法限制個別教師的影響範圍,結合多位教師會導致其能力之間出現依賴於混合比例的蹺蹺板效應。這些結果闡明了OPD何時能泛化,並為診斷多教師OPD提供了有用的視角。
English
On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.