ChatPaper.aiChatPaper

SPOT: オンポリシー蒸留のためのスパース探索と結果較正

SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation

August 5, 2026
著者: Zikun Qu, Min Zhang, Mingze Kong, Zhiwei Shang, Yikun Ban, Shuang Qiu, Zhongxiang Dai
cs.AI

要旨

オン・ポリシー蒸留(OPD)は、学生モデルが生成した軌道上で高密度な教師監督を提供するが、標準的な逆KL訓練では、他の妥当な継続候補に十分な確率を割り当てられないことがある。教師エントロピーだけでは、不確実性が少数の妥当な次トークンに集中しているのか、長い確率テールに分散しているのか、また学生がそれらの候補を既に適切に表現しているのかは判別できない。さらに、局所的な教師確率は下流タスクの成功を予測しない場合がある。本稿では、Sparse Probing and Outcome-calibrated Targets OPD(SPOT)を導入する。これは、「どこをプローブするか」と「何を蒸留するか」という2つの結合した決定を、獲得—探索—活用の手順によって扱う。獲得段階では、正規化された教師エントロピー、小さなtop-k候補集合が捕捉する確率質量、および学生と教師のミスマッチを組み合わせた位置レベルのスコアを用いて、限られたプローブ予算を配分する。探索段階では、SPOTは検証器によってスコア付けされた学生の継続生成を通じて、教師が提案する候補を評価する。活用段階では、これらの結果から閉形式のKL正規化ターゲットが得られ、下流の成果がより良い候補を優先しつつ、教師分布に固定されたままでいる。複数の学生モデルと推論ベンチマークにわたる広範な実験により、解の品質とカバレッジのバランスを取りながら推論性能を向上させるSPOTの有効性が実証される。
English
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other plausible continuations. Teacher entropy alone does not reveal whether uncertainty is concentrated among a few plausible next tokens or dispersed over a long probability tail, nor whether the student already represents those candidates well. Moreover, local teacher probabilities may not predict downstream success. We introduce Sparse Probing and Outcome-calibrated Targets OPD (SPOT), which addresses two coupled decisions, where to probe and what to distill, through an acquisition--exploration--exploitation procedure. During acquisition, a position-level score combines normalized teacher entropy, the probability mass captured by a small top-k candidate set, and student--teacher mismatch to allocate a limited probing budget. During exploration, SPOT evaluates teacher-proposed candidates through verifier-scored student continuations. During exploitation, these outcomes produce a closed-form, KL-regularized target that favors candidates with better downstream outcomes while remaining anchored to the teacher distribution. Extensive experiments across multiple student models and reasoning benchmarks demonstrate the effectiveness of SPOT in improving reasoning performance while balancing solution quality and coverage.