Flux-OPD:基于演化上下文的在线策略蒸馏
Flux-OPD: On-Policy Distillation with Evolving Contexts
July 30, 2026
作者: Yuran Wang, Zekun Wang, Bohan Zeng, Ruixu Zhang, Wenxuan Liu, Liu Yang, Yifan Dai, Yang Shi, Bozhou Li, Chengzhuo Tong, Daili Hua, Yuanxing Zhang, Wentao Zhang
cs.AI
摘要
在开放式领域中,大语言模型的训练缺乏可验证的奖励,使得任务偏好难以形式化为有效的监督信号。上下文可以传达此类偏好,但一旦被蒸馏到学生模型中,其提供的额外监督便十分有限,这促使人们探索随学生表现演化的上下文。然而,直接将演化中的上下文用作训练过程中的监督,会导致蒸馏目标不稳定以及分布相互冲突,因此需要相应的机制来稳定目标并降低冲突的权重。本文通过对反向KL目标进行分解来分析上下文的影响,并揭示了两个发现:学生模型被向着上下文条件化教师模型的几何平均进行蒸馏;同时,该目标包含一个冲突项,用于衡量这些教师模型之间的冲突程度。基于这一分解,我们提出Flux-OPD,一种在线策略蒸馏(OPD)范式,利用演化中的上下文作为训练过程中的监督,以捕捉开放式领域中的任务偏好。Flux-OPD将上下文条件化教师与无上下文教师之间的差异视为上下文差异信号,将其作为上下文修正注入无上下文教师锚点,并以冲突项为指标对这些修正的强度进行加权。在开放式任务上的实验表明,Flux-OPD优于现有的OPD范式,凸显了将教师监督与演化上下文相结合的潜力。
English
Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and conflicting distributions, requiring mechanisms to stabilize target and downweight conflicts. In this paper, we analyze the effect of contexts through a decomposition of the reverse KL objective, revealing two findings: the student is distilled toward the geometric mean of context-conditioned teachers, and the objective contains a conflict term that measures conflicts among these teachers. Based on this decomposition, we propose Flux-OPD, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains. Flux-OPD treats the differences between context-conditioned and context-free teachers as contextual difference signals, injects them as contextual corrections into the context-free teacher anchor, and weights their correction strength using the conflict term as an indicator. Experiments on open-ended tasks show that Flux-OPD outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.