ChatPaper.aiChatPaper

オン方策拡散蒸留における分類器不要ガイダンスの再考

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

July 27, 2026
著者: Bingnan Li, Haozhe Wang, Haozhong Xiong, Fangtai Wu, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang
cs.AI

要旨

オン・ポリシー蒸留(OPD)は、現在の生徒モデルが生成する軌跡に沿って教師モデルを照会することにより拡散モデルを適応させる手法であるが、現代の拡散システムのデフォルト構成要素である分類器不要ガイダンス(CFG)の下でどのように動作すべきかについては、未解明の部分が多い。既存のOPD手法は、速度マッチングをCFG合成予測に自然に拡張し、教師と生徒のガイド付き速度を直接一致させる。我々は、この目的関数がブランチレベルで未同定性を持つことを示す。すなわち、正ブランチ誤差と負ブランチ誤差がガイド付き予測において相殺し合う可能性がある。二つの対照的なケースを通じて、単純なマッチングが共有負条件付けの下では有効であり、両ブランチの誤差が共同で減少することを見出した。しかし、モデルの本来のCFGスキームが教師の負ブランチに生徒が利用できない特権情報を保持している場合、この共同減少は崩壊し、合成目的関数は敵対的なブランチ誤差ダイナミクスを誘発し、正ブランチ誤差を減少させる一方で負ブランチ誤差を増大させる。我々はこの故障モードをNegative Branch Asymmetry(NBA)と名付ける。NBAに対処するため、正予測とCFG条件方向を別々に制約するブランチ認識型OPD目的関数であるPositive-Direction Matching(PDM)を導入する。我々はPDMを密から疎へのビデオ制御に適用する。このタスクでは、単純なガイド付きマッチングが推論ガイダンス尺度に非常に敏感である一方、ブランチ認識型の監督はより頑健で効果的な知識伝達を可能にする。
English
On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the CFG-composed prediction, directly matching teacher and student guided velocities. We show that this objective is under-identified at the branch level: positive- and negative-branch errors can compensate in the guided prediction. Through two contrasting cases, we find that naive matching remains effective under shared negative conditioning, where both branch errors decrease jointly. When the model's native CFG schema retains privileged information in the teacher's negative branch that is unavailable to the student, however, this joint reduction breaks down and the composed objective induces antagonistic branch-error dynamics, reducing the positive-branch error while increasing the negative-branch error. We term this failure mode Negative Branch Asymmetry (NBA). To address NBA, we introduce Positive--Direction Matching (PDM), a branch-aware OPD objective that separately constrains the positive prediction and the CFG conditional direction. We apply PDM to dense-to-sparse video control, where naive guided matching is highly sensitive to inference guidance scales, while branch-aware supervision enables more robust and effective knowledge transfer.