拡散モデルにおけるオン方策自己蒸留
On-Policy Self-Distillation in Diffusion Models
August 25, 2026
著者: Wei Zhou, Xiongwei Zhu, Lingdong Kong, Bo Chen, Lei Zhang, Yongyuan Liang, Xiaoxia Hou, Ye Tian, Xian Sun, Yingshuo Wang, Linfeng Li, Shengqiong Wu, Leigang Qu, Feng Li, Wei Liu, Julian McAuley, Tat-Seng Chua
cs.AI
要旨
強化学習は拡散モデルを人間の選好やタスク固有の目的に整合させることができるが、終点報酬は中間のノイズ除去予測がどのように変化すべきかを指定しない。我々は、画像レベルの報酬ガイダンスを、サンプリングされたクエリにおけるクリーン出力予測の明示的なターゲットに変換する、オンポリシー自己蒸留フレームワークであるDiffusionOPSDを導入する。各外部反復において、凍結された行動ポリシーが軌跡を生成し、クエリ状態とアンカーを提供する。報酬勾配は、各アンカーの周りに有界な正および負のターゲットを構築する。学習可能なポリシーは、指数移動平均更新が行動ポリシーを更新する前に、有限回のフィッティングを通じて、これらのターゲットを切り離された教師信号として適合させる。この設定により、ターゲット構築と有限実現を別々に測定できる。制御された同一クエリ実験では、ターゲット構築の利得が大きいほど、単一のフィッティング更新後の実現利得が大きくなるとは限らないことが示される。SD 3.5-Mとステップ蒸留されたZ-Image-Turboにおいて、我々のアプローチは、2つのバックボーンと10の評価器にわたる20の報酬整合設定のうち19で、最終的なホールドアウトスコアの最良値を達成する。これは最強の競合手法を最大44.0%上回り、訓練GPU時間をDiffusionNFTと比較してSD 3.5-Mでは40%、Z-Image-Turboでは63%削減する。これらの結果は、画像レベルの報酬ガイダンスを明示的かつ継続的に更新される中間教師信号に変換することにより、オンポリシー自己蒸留が拡散モデルのポストトレーニングに対する効率的かつ分析可能なアプローチであることを支持し、より効率的で診断可能なアライメントへの道を開く。
English
Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint rewards do not specify how an intermediate denoising prediction should change. We introduce DiffusionOPSD as an on-policy self-distillation framework that converts image-level reward guidance into explicit targets for clean-output predictions at sampled queries. At each outer iteration, a frozen behavior policy generates trajectories and supplies query states and anchors. Reward gradients construct bounded positive and negative targets around each anchor. The trainable policy fits these targets as detached supervision through finite fitting before an exponential moving average update refreshes the behavior policy. This setup lets us measure target construction and finite realization separately. Controlled same-query experiments show that larger target-construction gains do not necessarily translate into larger realized gains after a single fitting update. Across SD 3.5-M and the step-distilled Z-Image-Turbo, our approach achieves the best final held-out scores in 19 of 20 reward-matched settings across two backbones and ten evaluators. It outperforms the strongest competing method by up to 44.0% and reduces training GPU-hours relative to DiffusionNFT by 40% on SD 3.5-M and 63% on Z-Image-Turbo. These results support on-policy self-distillation as an efficient and analyzable approach to diffusion post-training by converting image-level reward guidance into explicit and continually refreshed intermediate supervision, thereby opening a path toward more efficient and diagnosable alignment.