ChatPaper.aiChatPaper

온-정책 확산 증류에서의 분류기 프리 안내 재고

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

July 27, 2026
저자: Bingnan Li, Haozhe Wang, Haozhong Xiong, Fangtai Wu, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang
cs.AI

초록

온-정책 증류(OPD)는 현재 학생이 생성한 궤적을 따라 교사 모델에 질의하여 확산 모델을 적응시키지만, 현대 확산 시스템의 기본 구성 요소인 분류기-없는 유도(CFG) 하에서 어떻게 작동해야 하는지는 아직 충분히 이해되지 않았다. 기존 OPD 방법은 속도 정합을 CFG 구성 예측으로 자연스럽게 확장하여 교사와 학생의 유도 속도를 직접 정합한다. 우리는 이 목적함수가 가지 수준에서 저식별됨을 보인다: 양성 및 음성 가지 오차가 유도 예측에서 상쇄될 수 있다. 두 가지 대조 사례를 통해, 순진한 정합이 공유된 음성 조건화 하에서는 효과적임을 발견한다. 여기서는 두 가지 오차가 함께 감소한다. 그러나 모델의 고유 CFG 스키마가 교사의 음성 가지에 학생이 접근할 수 없는 특권 정보를 유지할 때, 이러한 공동 감소는 무너지고, 구성된 목적함수는 양성 가지 오차를 줄이면서 음성 가지 오차를 증가시키는 대립적 가지-오차 동역학을 유도한다. 우리는 이 실패 모드를 음성 가지 비대칭(NBA)이라 명명한다. NBA를 해결하기 위해, 우리는 양성-방향 정합(PDM)을 소개한다. 이는 양성 예측과 CFG 조건부 방향을 별도로 제약하는 가지 인식 OPD 목적함수이다. 우리는 PDM을 밀집-희소 비디오 제어에 적용한다. 여기서 순진한 유도 정합은 추론 유도 척도에 매우 민감한 반면, 가지 인식 지도 학습은 더 강건하고 효과적인 지식 전이를 가능하게 한다.
English
On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the CFG-composed prediction, directly matching teacher and student guided velocities. We show that this objective is under-identified at the branch level: positive- and negative-branch errors can compensate in the guided prediction. Through two contrasting cases, we find that naive matching remains effective under shared negative conditioning, where both branch errors decrease jointly. When the model's native CFG schema retains privileged information in the teacher's negative branch that is unavailable to the student, however, this joint reduction breaks down and the composed objective induces antagonistic branch-error dynamics, reducing the positive-branch error while increasing the negative-branch error. We term this failure mode Negative Branch Asymmetry (NBA). To address NBA, we introduce Positive--Direction Matching (PDM), a branch-aware OPD objective that separately constrains the positive prediction and the CFG conditional direction. We apply PDM to dense-to-sparse video control, where naive guided matching is highly sensitive to inference guidance scales, while branch-aware supervision enables more robust and effective knowledge transfer.