Flux-OPD: 진화하는 맥락을 활용한 온-정책 증류
Flux-OPD: On-Policy Distillation with Evolving Contexts
July 30, 2026
저자: Yuran Wang, Zekun Wang, Bohan Zeng, Ruixu Zhang, Wenxuan Liu, Liu Yang, Yifan Dai, Yang Shi, Bozhou Li, Chengzhuo Tong, Daili Hua, Yuanxing Zhang, Wentao Zhang
cs.AI
초록
개방형 도메인에서의 대규모 언어 모델 훈련은 검증 가능한 보상이 부족하여 작업 선호도를 효과적인 감독으로 공식화하기 어렵다. 컨텍스트는 그러한 선호도를 전달할 수 있지만, 학생 모델로 증류된 이후에는 추가적인 감독을 거의 제공하지 못하므로, 학생의 성과에 맞춰 진화하는 컨텍스트가 필요하다. 그러나 진화하는 컨텍스트를 훈련 중 감독으로 직접 사용하면 증류 대상이 불안정해지고 분포 간 충돌이 발생하므로, 대상을 안정화하고 충돌의 영향력을 낮추는 메커니즘이 요구된다. 본 논문에서는 역방향 KL 목적 함수의 분해를 통해 컨텍스트의 효과를 분석하여 두 가지 결과를 밝힌다. 첫째, 학생 모델은 컨텍스트 조건부 교사들의 기하 평균으로 증류된다. 둘째, 목적 함수에는 이러한 교사들 간의 충돌을 측정하는 충돌 항이 포함된다. 이러한 분해를 바탕으로, 본 논문은 개방형 도메인에서 작업 선호도를 포착하기 위해 진화하는 컨텍스트를 훈련 중 감독으로 사용하는 OPD 패러다임인 Flux-OPD를 제안한다. Flux-OPD는 컨텍스트 조건부 교사와 컨텍스트 비조건부 교사 간의 차이를 컨텍스트 차이 신호로 취급하고, 이를 컨텍스트 비조건부 교사 앵커에 컨텍스트 보정으로 주입하며, 충돌 항을 지표로 사용하여 보정 강도에 가중치를 부여한다. 개방형 작업에 대한 실험 결과, Flux-OPD는 기존 OPD 패러다임보다 우수한 성능을 보였으며, 이는 교사 감독과 진화하는 컨텍스트를 결합할 수 있는 잠재력을 강조한다.
English
Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and conflicting distributions, requiring mechanisms to stabilize target and downweight conflicts. In this paper, we analyze the effect of contexts through a decomposition of the reverse KL objective, revealing two findings: the student is distilled toward the geometric mean of context-conditioned teachers, and the objective contains a conflict term that measures conflicts among these teachers. Based on this decomposition, we propose Flux-OPD, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains. Flux-OPD treats the differences between context-conditioned and context-free teachers as contextual difference signals, injects them as contextual corrections into the context-free teacher anchor, and weights their correction strength using the conflict term as an indicator. Experiments on open-ended tasks show that Flux-OPD outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.