論擴散模型中的同策略自蒸餾
On-Policy Self-Distillation in Diffusion Models
August 25, 2026
作者: Wei Zhou, Xiongwei Zhu, Lingdong Kong, Bo Chen, Lei Zhang, Yongyuan Liang, Xiaoxia Hou, Ye Tian, Xian Sun, Yingshuo Wang, Linfeng Li, Shengqiong Wu, Leigang Qu, Feng Li, Wei Liu, Julian McAuley, Tat-Seng Chua
cs.AI
摘要
強化學習能將擴散模型與人類偏好及任務特定目標對齊,但端點獎勵並未指明中間去噪預測應如何改變。我們提出DiffusionOPSD作為一個在策略自我蒸餾框架,將影像層級的獎勵引導轉換為在取樣查詢點上對乾淨輸出預測的明確目標。在每次外部迭代中,一個凍結的行為策略生成軌跡,並提供查詢狀態與錨點。獎勵梯度在每個錨點周圍建構有界的正目標與負目標。可訓練策略透過有限次擬合將這些目標作為分離監督來擬合,隨後指數移動平均更新刷新行為策略。此設置使我們能夠分別衡量目標建構與有限實現的效果。受控的同查詢實驗顯示,在單次擬合更新後,較大的目標建構增益不一定轉化為較大的實現增益。在SD 3.5-M與步進蒸餾的Z-Image-Turbo上,我們的方法在兩個骨幹網路與十個評估器的20組獎勵匹配設定中,於19組取得最佳最終留出分數。其表現優於最強競爭方法最多達44.0%,且相較於DiffusionNFT,訓練GPU小時數在SD 3.5-M上減少40%,在Z-Image-Turbo上減少63%。這些結果支持在策略自我蒸餾作為一種高效且可分析的擴散模型後訓練方法,透過將影像層級獎勵引導轉換為明確且持續刷新的中間監督,從而為更高效且可診斷的對齊開闢路徑。
English
Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint rewards do not specify how an intermediate denoising prediction should change. We introduce DiffusionOPSD as an on-policy self-distillation framework that converts image-level reward guidance into explicit targets for clean-output predictions at sampled queries. At each outer iteration, a frozen behavior policy generates trajectories and supplies query states and anchors. Reward gradients construct bounded positive and negative targets around each anchor. The trainable policy fits these targets as detached supervision through finite fitting before an exponential moving average update refreshes the behavior policy. This setup lets us measure target construction and finite realization separately. Controlled same-query experiments show that larger target-construction gains do not necessarily translate into larger realized gains after a single fitting update. Across SD 3.5-M and the step-distilled Z-Image-Turbo, our approach achieves the best final held-out scores in 19 of 20 reward-matched settings across two backbones and ten evaluators. It outperforms the strongest competing method by up to 44.0% and reduces training GPU-hours relative to DiffusionNFT by 40% on SD 3.5-M and 63% on Z-Image-Turbo. These results support on-policy self-distillation as an efficient and analyzable approach to diffusion post-training by converting image-level reward guidance into explicit and continually refreshed intermediate supervision, thereby opening a path toward more efficient and diagnosable alignment.