ChatPaper.aiChatPaper

Flux-OPD: 進化するコンテキストによるオン方策蒸留

Flux-OPD: On-Policy Distillation with Evolving Contexts

July 30, 2026
著者: Yuran Wang, Zekun Wang, Bohan Zeng, Ruixu Zhang, Wenxuan Liu, Liu Yang, Yifan Dai, Yang Shi, Bozhou Li, Chengzhuo Tong, Daili Hua, Yuanxing Zhang, Wentao Zhang
cs.AI

要旨

オープンエンド領域における大規模言語モデルの学習は検証可能な報酬を欠いており、タスク選好を効果的な監督として形式化することを困難にしている。コンテキストはそのような選好を伝えることができるものの、生徒に蒸留されると追加の監督をほとんど提供しないため、生徒の性能に応じて進化するコンテキストが動機付けられる。しかしながら、進化するコンテキストを訓練中の監督として直接使用すると、蒸留ターゲットが不安定になり、分布が矛盾するため、ターゲットを安定化し、矛盾を低重み付けするメカニズムが必要となる。本論文では、逆KL目的関数の分解を通じてコンテキストの影響を分析し、以下の2つの知見を得る。すなわち、生徒はコンテキスト条件付き教師の幾何平均へ向けて蒸留され、また目的関数にはこれらの教師間の矛盾を測定する矛盾項が含まれる。この分解に基づき、オープンエンド領域におけるタスク選好を捉えるために、進化するコンテキストを訓練中の監督として使用するOPDパラダイムであるFlux-OPDを提案する。Flux-OPDは、コンテキスト条件付き教師とコンテキスト非依存教師の差分をコンテキスト差分シグナルとして扱い、それらをコンテキスト修正としてコンテキスト非依存教師アンカーに注入し、矛盾項を指標としてそれらの修正強度を重み付けする。オープンエンドタスクに関する実験により、Flux-OPDが既存のOPDパラダイムを上回ることが示され、教師監督と進化するコンテキストを組み合わせる可能性が強調される。
English
Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and conflicting distributions, requiring mechanisms to stabilize target and downweight conflicts. In this paper, we analyze the effect of contexts through a decomposition of the reverse KL objective, revealing two findings: the student is distilled toward the geometric mean of context-conditioned teachers, and the objective contains a conflict term that measures conflicts among these teachers. Based on this decomposition, we propose Flux-OPD, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains. Flux-OPD treats the differences between context-conditioned and context-free teachers as contextual difference signals, injects them as contextual corrections into the context-free teacher anchor, and weights their correction strength using the conflict term as an indicator. Experiments on open-ended tasks show that Flux-OPD outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.