重新思考在策略扩散蒸馏中的无分类器引导
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
July 27, 2026
作者: Bingnan Li, Haozhe Wang, Haozhong Xiong, Fangtai Wu, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang
cs.AI
摘要
在策略蒸馏(OPD)通过沿着当前学生模型生成的轨迹查询教师模型来适配扩散模型,但在无分类器引导(CFG,现代扩散系统的默认组件)下其行为模式尚不明确。现有OPD方法自然地将速度匹配扩展至CFG组合预测,直接对齐教师与学生模型的引导速度。我们证明该目标在分支层面存在欠辨识问题:正负分支误差可在引导预测中相互补偿。通过两个对比案例发现,在共享负条件下(即两分支误差协同下降),朴素匹配方法仍然有效。然而,当模型原生CFG架构在教师负分支中保留学生无法获取的特权信息时,这种协同下降机制失效,组合目标会引发对抗性分支误差动力学:在降低正分支误差的同时增加负分支误差。我们将此失效模式称为负分支非对称性(NBA)。为应对NBA,我们提出正向方向匹配(PDM),这是一种分支感知的OPD目标,可独立约束正向预测与CFG条件方向。我们将PDM应用于稠密到稀疏视频控制任务,在此场景下,朴素引导匹配对推理引导尺度高度敏感,而分支感知监督能实现更鲁棒有效的知识迁移。
English
On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the CFG-composed prediction, directly matching teacher and student guided velocities. We show that this objective is under-identified at the branch level: positive- and negative-branch errors can compensate in the guided prediction. Through two contrasting cases, we find that naive matching remains effective under shared negative conditioning, where both branch errors decrease jointly. When the model's native CFG schema retains privileged information in the teacher's negative branch that is unavailable to the student, however, this joint reduction breaks down and the composed objective induces antagonistic branch-error dynamics, reducing the positive-branch error while increasing the negative-branch error. We term this failure mode Negative Branch Asymmetry (NBA). To address NBA, we introduce Positive--Direction Matching (PDM), a branch-aware OPD objective that separately constrains the positive prediction and the CFG conditional direction. We apply PDM to dense-to-sparse video control, where naive guided matching is highly sensitive to inference guidance scales, while branch-aware supervision enables more robust and effective knowledge transfer.