ChatPaper.aiChatPaper

시각적 대조적 자기 증류

Visual Contrastive Self-Distillation

July 23, 2026
저자: Yijun Liang, Yunjie Tian, Yijiang Li, Yuqi Jia, Furong Huang, Tianyi Zhou, Di Fu
cs.AI

초록

온-정책 자기 증류(OPSD)는 온-정책 증류(OPD)에서 요구되는 외부 교사를 제거한다는 점에서 유망하지만, 교사가 학생보다 더 강력한 학습 신호를 제공하도록 보장하기 위해 여전히 교사와 학생 간의 비대칭 정보가 필요하다. 기존 방법들은 특권 답변 또는 시각적 증거를 통해 이러한 비대칭성을 창출한다. 본 논문에서는 이 두 요소를 모두 제거하여, 순전히 입력 조건화에 의해 구동되는 더 간단한 형태의 OPSD가 가능한지 질문한다. 이를 위해 우리는 시각적 대조 자기 증류(VCSD)를 제안하며, 이는 이미지 콘텐츠 제거를 온-정책 자기 증류 신호로 변환한다. 각 학생이 생성한 응답 프리픽스에서, EMA 교사는 동일한 프롬프트와 프리픽스 하에서 두 가지 다음 토큰 분포를 생성한다. 하나는 원본 이미지에 조건화된 분포이고, 다른 하나는 콘텐츠가 삭제된 제어 조건에 조건화된 분포이다. 이들의 토큰별 로그 확률 차이는 인스턴스 수준의 시각적 콘텐츠에 의해 우도가 구체적으로 증가된 후보들을 부각시킨다. 우리는 이 대비를 사용하여 교사의 원본 이미지 분포를 그 가능한 지지 영역 내에서 선명화하고, 결과적으로 얻은 전체 분포 목표를 학생에게 증류한다. ViRL39K 데이터셋을 사용한 결과, VCSD는 Qwen3-VL 및 Qwen3.5 모델 전반에 걸쳐 대응되는 OPSD보다 일관되게 우수한 성능을 보였다. 예를 들어, Qwen3-VL에서 2B는 62.27%에서 67.04%로, 4B는 71.30%에서 73.16%로, 8B는 72.51%에서 76.26%로 7개 벤치마크 종합 점수가 개선되었다. 또한 VCSD는 외부 교사, 특권 답변, 시각적 증거 신호, 추론 궤적 또는 추가적인 추론 시간 비용이 필요하지 않다.
English
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods create this asymmetry either through privileged answers or visual evidence. We ask whether both can be removed, yielding a simpler form of OPSD driven purely by input conditioning. For this purpose, we propose Visual Contrastive Self-Distillation, namely VCSD, which converts image-content removal into an on-policy self-distillation signal. At each student-generated response prefix, the EMA teacher produces two next-token distributions under the same prompt and prefix -- one conditioned on the original image and the other on a content-erased control. Their token-wise log-probability difference highlights candidates whose likelihood is specifically increased by the instance-level visual content. We use this contrast to sharpen the teacher's original-image distribution within its plausible support, and distill the resulting full-distribution target into the student. Using ViRL39K dataset, VCSD consistently outperforms matched OPSD across Qwen3-VL and Qwen3.5 models. For example, on Qwen3-VL, it improves the seven-benchmark aggregate from 62.27% rightarrow 67.04% at 2B, 71.30% rightarrow 73.16% at 4B, and 72.51% rightarrow 76.26% at 8B. Furthermore, VCSD requires no external teacher, privileged answers, visual evidence signals, reasoning traces, or additional inference-time cost.