확산 모델에서의 온-폴리시 자기 증류
On-Policy Self-Distillation in Diffusion Models
August 25, 2026
저자: Wei Zhou, Xiongwei Zhu, Lingdong Kong, Bo Chen, Lei Zhang, Yongyuan Liang, Xiaoxia Hou, Ye Tian, Xian Sun, Yingshuo Wang, Linfeng Li, Shengqiong Wu, Leigang Qu, Feng Li, Wei Liu, Julian McAuley, Tat-Seng Chua
cs.AI
초록
강화 학습은 확산 모델을 인간의 선호와 작업별 목표에 정렬할 수 있지만, 종점 보상은 중간 단계의 노이즈 제거 예측이 어떻게 변화해야 하는지를 지정하지 않는다. 본 논문에서는 DiffusionOPSD를 온-정책 자기 증류 프레임워크로 소개한다. 이 프레임워크는 이미지 수준의 보상 안내를 샘플링된 쿼리에서 깨끗한 출력 예측을 위한 명시적 목표값으로 변환한다. 각 외부 반복에서 고정된 행동 정책이 궤적을 생성하고 쿼리 상태와 앵커를 제공한다. 보상 그래디언트는 각 앵커 주위에 유계의 양성 및 음성 목표값을 구성한다. 학습 가능한 정책은 지수 이동 평균 업데이트가 행동 정책을 갱신하기 전에 유한 피팅을 통해 이러한 목표값을 분리된 지도 학습으로 피팅한다. 이러한 설정을 통해 목표값 구성과 유한 실현을 별도로 측정할 수 있다. 동일 쿼리 통제 실험은 더 큰 목표값 구성 이득이 단일 피팅 업데이트 후 더 큰 실현 이득으로 반드시 이어지지 않음을 보여준다. SD 3.5-M과 단계 증류된 Z-Image-Turbo 전반에 걸쳐, 본 접근법은 두 백본과 열 개의 평가자에 걸친 20개의 보상 일치 설정 중 19개에서 최고의 최종 홀드아웃 점수를 달성했다. 가장 강력한 경쟁 방법보다 최대 44.0% 더 우수했으며, DiffusionNFT 대비 훈련 GPU 시간을 SD 3.5-M에서 40%, Z-Image-Turbo에서 63% 절감했다. 이러한 결과는 온-정책 자기 증류가 이미지 수준의 보상 안내를 명시적이고 지속적으로 갱신되는 중간 지도 학습으로 변환함으로써 확산 모델 사후 훈련에 대한 효율적이고 분석 가능한 접근법임을 뒷받침하며, 더 효율적이고 진단 가능한 정렬을 위한 경로를 연다.
English
Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint rewards do not specify how an intermediate denoising prediction should change. We introduce DiffusionOPSD as an on-policy self-distillation framework that converts image-level reward guidance into explicit targets for clean-output predictions at sampled queries. At each outer iteration, a frozen behavior policy generates trajectories and supplies query states and anchors. Reward gradients construct bounded positive and negative targets around each anchor. The trainable policy fits these targets as detached supervision through finite fitting before an exponential moving average update refreshes the behavior policy. This setup lets us measure target construction and finite realization separately. Controlled same-query experiments show that larger target-construction gains do not necessarily translate into larger realized gains after a single fitting update. Across SD 3.5-M and the step-distilled Z-Image-Turbo, our approach achieves the best final held-out scores in 19 of 20 reward-matched settings across two backbones and ten evaluators. It outperforms the strongest competing method by up to 44.0% and reduces training GPU-hours relative to DiffusionNFT by 40% on SD 3.5-M and 63% on Z-Image-Turbo. These results support on-policy self-distillation as an efficient and analyzable approach to diffusion post-training by converting image-level reward guidance into explicit and continually refreshed intermediate supervision, thereby opening a path toward more efficient and diagnosable alignment.