Flux-OPD:動態情境下的在線策略蒸餾
Flux-OPD: On-Policy Distillation with Evolving Contexts
July 30, 2026
作者: Yuran Wang, Zekun Wang, Bohan Zeng, Ruixu Zhang, Wenxuan Liu, Liu Yang, Yifan Dai, Yang Shi, Bozhou Li, Chengzhuo Tong, Daili Hua, Yuanxing Zhang, Wentao Zhang
cs.AI
摘要
在開放式領域中,大型語言模型的訓練缺乏可驗證的回報,使得任務偏好難以被形式化為有效的監督。上下文可以傳達這類偏好,然而一旦被蒸餾至學生模型中,便只能提供很少的額外監督,這激勵了隨學生表現而演化的上下文。然而,直接將演化中的上下文用作訓練中的監督,會導致不穩定的蒸餾目標與相互衝突的分布,因此需要機制來穩定目標並降低衝突的權重。在本文中,我們透過對反向KL目標的分解來分析上下文的效果,揭示了兩個發現:學生會被蒸餾至上下文條件教師的幾何平均,且該目標包含一個衡量這些教師之間衝突的衝突項。基於此分解,我們提出了Flux-OPD,一種使用演化上下文作為訓練中監督、以捕捉開放式領域任務偏好的OPD範式。Flux-OPD將上下文條件教師與無上下文教師之間的差異視為上下文差異信號,將其作為上下文校正注入到無上下文教師錨點中,並以衝突項作為指標來加權其校正強度。在開放式任務上的實驗表明,Flux-OPD優於現有的OPD範式,凸顯了將教師監督與演化上下文結合的潛力。
English
Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and conflicting distributions, requiring mechanisms to stabilize target and downweight conflicts. In this paper, we analyze the effect of contexts through a decomposition of the reverse KL objective, revealing two findings: the student is distilled toward the geometric mean of context-conditioned teachers, and the objective contains a conflict term that measures conflicts among these teachers. Based on this decomposition, we propose Flux-OPD, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains. Flux-OPD treats the differences between context-conditioned and context-free teachers as contextual difference signals, injects them as contextual corrections into the context-free teacher anchor, and weights their correction strength using the conflict term as an indicator. Experiments on open-ended tasks show that Flux-OPD outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.