ChatPaper.aiChatPaper

一つの症状、三つの梃子:オンポリシー自己蒸留の批判的レビュー

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

August 26, 2026
著者: Justin Robert, Raheel Qader
cs.AI

要旨

オンポリシー蒸留は、言語モデルを自身の生成出力に対して訓練し、教師がそれらをトークン単位で採点する手法である。これは、模倣学習の密な教師信号と、強化学習のオンポリシーサンプリングを組み合わせたものである。しかし、この手法には教師役となる、より大きな第二のモデルが必要である。オンポリシー自己蒸留(OPSD)は、そのコストを取り除く。OPSDでは、教師はモデル自身であり、テスト時に生徒が利用できない特権情報、例えば参照解、プラン、環境フィードバックなどを条件として与えられる。教師は生徒より強力なわけではなく、より多くの情報を与えられているだけである。初期の結果は有望であり、強化学習に匹敵する精度を、大幅に少ない生成トークン数で達成した。しかし、シグナルを生み出すその非対称性は、同時にシグナルにバイアスを与える。現在、この分野で支配的な失敗モードは崩壊(collapse)、すなわちモデルが生成可能な推論経路の集合が徐々に狭まる現象である。崩壊はOPSDに固有のものではないが、特権情報がそれを悪化させる。本レビューでは、崩壊を三つの操作因子によって支配される症状として扱う。すなわち、(i)シグナルが適用される場所、つまりトークンの重み付けの仕方、(ii)教師に示される内容、つまり特権情報の性質、(iii)シグナルが変化する時期、つまり教師のダイナミクスとガイダンスの減衰、である。本レビューは、本手法が生まれた領域であり、その失敗モードが最もよく記録されている数学的推論に対象を限定する。新たな実験の報告はない。その貢献は構造的なものである。すなわち、論文ごとに異なる名前で呼ばれてきた現象に共通の用語を提供し、確定している点と未だ議論のある点とを明確に区別することである。
English
On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Distillation (OPSD) removes that cost. The teacher is the model itself, conditioned on privileged information the student will not have at test time, such as a reference solution, a plan, or environment feedback. The teacher is no stronger than the student, only better informed. Early results were promising, with accuracy comparable to reinforcement learning at a fraction of the generated tokens. But the same asymmetry that produces the signal also biases it. One failure mode now dominates the field: collapse, the progressive narrowing of the set of reasoning paths the model can produce. Collapse is not specific to OPSD, though privileged information aggravates it. This review treats collapse as a symptom governed by three levers: (i) where the signal is applied, that is, how tokens are weighted; (ii) what the teacher is shown, that is, the nature of the privileged information; and (iii) when the signal changes, that is, the teacher's dynamics and the decay of guidance. We restrict our scope to mathematical reasoning, where the method originated and where its failure modes are best documented. We report no new experiments. The contribution is structural: a shared vocabulary for phenomena named differently across papers, and a clear line between what is settled and what is still disputed.