ChatPaper.aiChatPaper

重新審視同策略擴散蒸餾中的無分類器引導

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

July 27, 2026
作者: Bingnan Li, Haozhe Wang, Haozhong Xiong, Fangtai Wu, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang
cs.AI

摘要

同策略蒸餾(OPD)透過沿當前學生模型生成的軌跡查詢教師模型來調整擴散模型,但其在無分類器引導(CFG)——現代擴散系統的預設組件——下的運作方式仍未被充分理解。現有OPD方法自然將速度匹配延伸至由CFG組合成的預測,直接比對教師與學生模型的引導速度。我們證明此目標在分支層級存在欠識別問題:正分支與負分支誤差可在引導預測中相互補償。透過兩個對比案例,我們發現樸素匹配在共享負條件下仍有效,此時兩個分支誤差共同下降。然而,當模型原生CFG架構在教師模型的負分支中保留學生無法取得的特權資訊時,此共同下降機制便失效,組合目標將引發對抗性的分支誤差動態——降低正分支誤差卻增加負分支誤差。我們將此失敗模式稱為「負分支不對稱性」(NBA)。為解決NBA,我們提出「正向方向匹配」(PDM),這是一種分支感知的OPD目標,能分別約束正向預測與CFG條件方向。我們將PDM應用於密集至稀疏影片控制,在此場景中,樸素引導匹配對推論引導尺度非常敏感,而分支感知監督則能實現更穩健有效的知識遷移。
English
On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the CFG-composed prediction, directly matching teacher and student guided velocities. We show that this objective is under-identified at the branch level: positive- and negative-branch errors can compensate in the guided prediction. Through two contrasting cases, we find that naive matching remains effective under shared negative conditioning, where both branch errors decrease jointly. When the model's native CFG schema retains privileged information in the teacher's negative branch that is unavailable to the student, however, this joint reduction breaks down and the composed objective induces antagonistic branch-error dynamics, reducing the positive-branch error while increasing the negative-branch error. We term this failure mode Negative Branch Asymmetry (NBA). To address NBA, we introduce Positive--Direction Matching (PDM), a branch-aware OPD objective that separately constrains the positive prediction and the CFG conditional direction. We apply PDM to dense-to-sparse video control, where naive guided matching is highly sensitive to inference guidance scales, while branch-aware supervision enables more robust and effective knowledge transfer.