ChatPaper.aiChatPaper

중간 학습 중 지식 증류는 사실 회상보다 추론을 선호한다

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

September 1, 2026
저자: Jacqueline He, Howard Yen, Shuyue Stella Li, Margaret Li, Hanqing Zeng, Yinglong Xia, Benyu Zhang, Zhuokai Zhao, Qiang Zhang, Pang Wei Koh, Luke Zettlemoyer, Wen-tau Yih
cs.AI

초록

로짓 기반 지식 증류(KD)는 더 강력한 교사 모델의 지도를 통해 더 작은 언어 모델(LM)을 훈련시키는 데 사용된다. 그러나 그러한 이점이 훈련 단계 전반에 걸쳐 일관적인지 여부는 아직 명확하지 않다. 통제된 실험을 통해, 우리는 사후 훈련(post-training)된 교사로 수행하는 순방향 KL 증류(표준 KD 공식)가 큐레이션된 말뭉치에 대한 자기지도 학습의 중간 단계인 중간 훈련(mid-training) 동안 근본적으로 다른 양상을 보임을 발견한다. 놀랍게도 순방향 KD는 표준 다음 토큰 예측(NTP) 대비 사전 훈련 중에는 추론 능력과 사실 회상을 동시에 향상시키지만, 중간 훈련 중에는 추론 향상이 지속됨에도 불구하고 사실 회상 습득을 오히려 늦춘다. 우리는 이러한 단계 의존성이 데이터 도메인 간 교사 확신도의 비대칭성과 학생 모델의 진화하는 지식 상태에 기인함을 추적한다. 즉, 교사는 지식 집약적 데이터보다 절차적 데이터에 대해 더 확신하는 반면, 학생 모델은 훈련 초기에 낮은 엔트로피의 사실적 지식을 습득한다. 이러한 불균형을 완화하기 위해, 우리는 교사의 예측 엔트로피를 경량 라우팅 신호로 사용하여 교사가 확신하는 토큰에 대해서는 증류를 수행하고, 그 외의 경우에는 크로스 엔트로피로 대체하는 단순한 중간 훈련 목적 함수인 Switch Distillation을 제안한다. Switch Distillation은 교사 모델 크기와 관계없이 기존 증류 목적 함수를 일관되게 능가한다. 표준 NTP 대비, 추론 성능은 1.61~1.71배, 지식 및 상식 성능은 1.13~1.19배를 달성하면서 사실 회상의 96.7~96.8%를 보존한다. 결정적으로, 이러한 이점은 사후 훈련 이후에도 지속된다. Switch Distillation은 사실 회상 격차를 해소하면서도 추론과 지식 및 상식에서 각각 1.25~1.32배와 1.13~1.20배의 향상을 유지한다.
English
Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits are consistent across training stages remains unclear. Through controlled experiments, we find that forward Kullback-Leibler (KL) distillation--the standard KD formulation--with post-trained teachers behaves fundamentally differently during mid-training, an intermediate phase of self-supervised learning on curated corpora. Surprisingly, while forward KD simultaneously improves reasoning and factual recall during pre-training relative to standard next-token prediction (NTP), it instead slows factual recall acquisition during mid-training despite continued reasoning gains. We trace this stage dependence to an asymmetry in teacher confidence across data domains and the student's evolving knowledge state: teachers are more confident on procedural than knowledge-intensive data, while students acquire low-entropy factual knowledge earlier in training. To mitigate this imbalance, we propose Switch Distillation, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy. Switch Distillation consistently outperforms existing distillation objectives across teacher sizes. Relative to standard NTP, it achieves 1.61-1.71x the reasoning performance and 1.13-1.19x the knowledge and commonsense performance while preserving 96.7-96.8% of factual recall. Crucially, these benefits persist after post-training: Switch Distillation closes the factual recall gap while maintaining 1.25-1.32x and 1.13-1.20x gains in reasoning and knowledge and commonsense, respectively.