ChatPaper.aiChatPaper

D-OPSD: Auto-Destilação On-Policy para o Ajuste Contínuo de Modelos de Difusão Destilados por Etapas

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models

May 6, 2026
Autores: Dengyang Jiang, Xin Jin, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Ruoyi Du, Xiangpeng Yang, Qilong Wu, Zhen Li, Peng Gao, Harry Yang, Steven Hoi
cs.AI

Resumo

O cenário dos modelos de geração de imagens de alto desempenho está atualmente a mudar dos ineficientes modelos multi-etapa para as suas contrapartes eficientes de poucas etapas (por exemplo, Z-Image-Turbo e FLUX.2-klein). No entanto, estes modelos apresentam desafios significativos para o ajuste fino supervisionado diretamente contínuo. Por exemplo, a aplicação da técnica de ajuste fino comumente utilizada compromete a sua capacidade inerente de inferência em poucas etapas. Para resolver isto, propomos o D-OPSD, um novo paradigma de treino para modelos de difusão com destilação de etapas que permite a aprendizagem *on-policy* durante o ajuste fino supervisionado. Primeiro, descobrimos que o modelo de difusão moderno, no qual o LLM/VLM funciona como codificador, pode herdar as capacidades *in-context* do seu codificador. Isto permite-nos transformar o treino num processo de auto-destilação *on-policy*. Especificamente, durante o treino, fazemos com que o modelo atue simultaneamente como professor e aluno com contextos diferentes: o aluno é condicionado apenas pela característica de texto, enquanto o professor é condicionado pela característica multimodal do *prompt* de texto e da imagem alvo. O treino minimiza as duas distribuições previstas sobre as próprias *roll-outs* do aluno. Ao ser otimizado na trajetória própria do modelo e sob a sua própria supervisão, o D-OPSD permite ao modelo aprender novos conceitos, estilos, etc., sem sacrificar a capacidade original de poucas etapas.
English
The landscape of high-performance image generation models is currently shifting from the inefficient multi-step ones to the efficient few-step counterparts (e.g, Z-Image-Turbo and FLUX.2-klein). However, these models present significant challenges for directly continuous supervised fine-tuning. For example, applying the commonly used fine-tuning technique would compromises their inherent few-step inference capability. To address this, we propose D-OPSD, a novel training paradigm for step-distilled diffusion models that enables on-policy learning during supervised fine-tuning. We first find that the modern diffusion model where the LLM/VLM serves as the encoder can inherit its encoder's in-context capabilities. This enables us to make the training as an on-policy self-distillation process. Specifically, during training, we make the model acts as both the teacher and the student with different contexts, where the student is conditioned only on the text feature, while the teacher is conditioned on the multimodal feature of both the text prompt and the target image. Training minimizes the two predicted distributions over the student's own roll-outs. By optimized on the model's own trajectory and under it's own supervision, D-OPSD enables the model to learn new concept, style, etc. without sacrificing the original few-step capacity.