모든 동전에는 양면이 있다: 대규모 언어 모델의 온-폴리시 증류에서 일반화의 이중적 특성에 대하여
Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models
August 17, 2026
저자: Zhaoyi Li, Deyang Kong, Yuan Wei, Evan Yang, Ranran Shen, Mahardika Krisna Ihsani, Ming Yang, Wei Zhang, Chuan Hao, Jian Yang, Ran Tao, Bryan Dai, Shikun Zhang, Wei Ye, Ying Wei, Defu Lian
cs.AI
초록
온-폴리시 증류(OPD)는 학생 자신의 정책에서 샘플링된 궤적을 통해 교사 능력을 전이하지만, 대부분의 연구가 단일 도메인과 훈련 데이터에 가까운 벤치마크에서 OPD를 평가함에 따라 그 일반화 행동은 여전히 제대로 이해되지 못하고 있다. 본 연구는 도메인 내 분포 변화부터 교차 도메인 전이 및 다중 교사 설정에 이르기까지, 한 번에 하나의 일반화 요인을 변화시키는 통제 연구를 제시한다. 우리는 OPD가 특정 문제에 대한 교사의 답변이 아니라 교사의 추론 행동을 전이한다는 것을 발견했다. 즉, 훈련 난이도는 거의 중요하지 않으며, 교사가 전혀 풀지 못한 문제조차 유용하다. 전이는 교사와 학생 간의 출처 관계에 크게 의존한다. 동일 출처 쌍은 언어, 추론 범위, 나아가 다른 도메인에 걸쳐 학생을 교사에 근접하게 만들지만, 교차 출처 쌍은 대부분 훈련된 분포에 적합할 뿐이다. 이러한 광범위한 영향력은 양날의 검이다. 도메인 전문가에게 프롬프트를 라우팅해도 각 교사의 영향을 국한시킬 수 없으므로, 교사들을 결합하면 그들의 능력 사이에서 혼합 비율에 의존하는 시소 효과가 발생하기 때문이다. 이러한 결과는 OPD가 언제 일반화되는지를 명확히 하며, 다중 교사 OPD를 진단하는 유용한 관점을 제공한다.
English
On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.