ChatPaper.aiChatPaper

검증 후 증류하라: 온폴리시 증류를 위한 프롬프트 수준의 교사 게이팅

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

September 2, 2026
저자: Zhiwei Zhang, Zechen Sun, Fei Zhao, Kang Peng, Bin Liang, Huayu Deng, Yao Hu, Kam-Fai Wong, Mu Chuan
cs.AI

초록

온-폴리시 증류(OPD)는 동결된 교사로부터 학생 자신의 롤아웃에 대해 토큰 수준의 조밀한 지도를 제공함으로써 사후 훈련을 가속화한다. 바닐라 OPD는 프롬프트별로 교사가 신뢰할 수 있는지 확인하지 않고 이 지도를 모든 프롬프트에 균일하게 적용한다. 역방향 KL 발산은 모드를 추구하는 성질이 있기 때문에, 확신을 가지고 틀린 교사는 강력하지만 잘못된 업데이트를 유도할 수 있다. 엔트로피나 교사-학생 우도 일치도와 같은 분포적 대리 지표는 불확실성이나 일치도를 측정할 뿐 결과의 정확성을 직접 검증하지는 않는다. 우리는 교사 신뢰도가 밀집 지도를 허용하기 전에 프롬프트 수준에서 검증되어야 한다는 원칙에 기반하여 TGOPD(Teacher-Gated On-Policy Distillation)를 도입한다. TGOPD는 검증기가 점수를 매긴 소규모 교사 프로브 집합에서 신뢰도를 추정하고, 신뢰도 검사를 통과한 각 프롬프트는 전적으로 밀집 OPD로 보내며, 그렇지 않은 경우에는 검증기 기반 GRPO로 보낸다. 수학, 코드, 지시 따르기 분야의 4B 및 35B 학생 모델에서 TGOPD는 여섯 가지 단일 도메인 설정 모두에서 바닐라 OPD를 능가하고, 다중 도메인 훈련에서는 두 규모 모두에서 더 높은 7개 벤치마크 평균을 달성한다. TGOPD는 그렇지 않으면 유휴 상태가 되는 교사 용량을 신뢰도 추정에 활용함으로써 비동기 OPD에서 교사 측 계산 낭비를 줄이며, 측정된 4B 단일 도메인 실행에서 교사 노드 GPU 활용률을 9.8%에서 78.9%로 높인다.
English
On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without checking whether the teacher is reliable for each prompt. Because reverse KL is mode-seeking, a confidently wrong teacher can induce a strong yet misleading update. Distributional proxies, such as entropy or teacher-student likelihood agreement, measure uncertainty or agreement but do not directly verify outcome correctness. We introduce Teacher-Gated On-Policy Distillation (TGOPD), built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted. TGOPD estimates reliability from a small set of verifier-scored teacher probes and routes each prompt exclusively to dense OPD when the reliability check passes or to verifier-grounded GRPO otherwise. Across 4B and 35B students in mathematics, code, and instruction following, TGOPD outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training. By using otherwise-idle teacher capacity for reliability estimation, TGOPD also reduces teacher-side compute waste in asynchronous OPD, increasing teacher-node GPU utilization from 9.8% to 78.9% in the measured 4B single-domain run.