ChatPaper.aiChatPaper

SPOT: 온-폴리시 증류를 위한 희소 프로빙 및 결과 보정

SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation

August 5, 2026
저자: Zikun Qu, Min Zhang, Mingze Kong, Zhiwei Shang, Yikun Ban, Shuang Qiu, Zhongxiang Dai
cs.AI

초록

온-폴리시 증류(OPD)는 학생이 생성한 궤적에 대해 조밀한 교사 감독을 제공하지만, 표준 역방향 KL 훈련은 다른 그럴듯한 연속 생성에 충분한 확률을 할당하지 못할 수 있다. 교사 엔트로피만으로는 불확실성이 소수의 유력한 다음 토큰에 집중되어 있는지, 아니면 긴 확률 꼬리에 분산되어 있는지, 그리고 학생이 그러한 후보들을 이미 잘 표현하고 있는지를 알 수 없다. 더욱이 국소적 교사 확률은 다운스트림 성공을 예측하지 못할 수 있다. 우리는 획득-탐색-활용 절차를 통해 ‘어디를 탐사할지’와 ‘무엇을 증류할지’라는 두 가지 결합된 결정을 다루는 SPOT(Sparse Probing and Outcome-calibrated Targets OPD)을 소개한다. 획득 단계에서는 위치 수준 점수가 정규화된 교사 엔트로피, 작은 상위-k 후보 집합이 포착하는 확률 질량, 그리고 학생-교사 불일치를 결합하여 제한된 탐사 예산을 배분한다. 탐색 단계에서 SPOT은 검증기가 점수를 매긴 학생 연속 생성을 통해 교사가 제안한 후보들을 평가한다. 활용 단계에서는 이러한 결과들이 폐쇄형 KL 정규화 타깃을 생성하며, 이는 더 나은 다운스트림 결과를 가진 후보들을 선호하면서도 교사 분포에 고정된 상태를 유지한다. 여러 학생 모델과 추론 벤치마크에 걸친 광범위한 실험은 SPOT이 솔루션 품질과 적용 범위의 균형을 유지하면서 추론 성능을 향상시키는 데 효과적임을 입증한다.
English
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other plausible continuations. Teacher entropy alone does not reveal whether uncertainty is concentrated among a few plausible next tokens or dispersed over a long probability tail, nor whether the student already represents those candidates well. Moreover, local teacher probabilities may not predict downstream success. We introduce Sparse Probing and Outcome-calibrated Targets OPD (SPOT), which addresses two coupled decisions, where to probe and what to distill, through an acquisition--exploration--exploitation procedure. During acquisition, a position-level score combines normalized teacher entropy, the probability mass captured by a small top-k candidate set, and student--teacher mismatch to allocate a limited probing budget. During exploration, SPOT evaluates teacher-proposed candidates through verifier-scored student continuations. During exploitation, these outcomes produce a closed-form, KL-regularized target that favors candidates with better downstream outcomes while remaining anchored to the teacher distribution. Extensive experiments across multiple student models and reasoning benchmarks demonstrate the effectiveness of SPOT in improving reasoning performance while balancing solution quality and coverage.