ChatPaper.aiChatPaper

자기지도 시각적 정책 기반 증류

Self-Supervised Visual On-Policy Distillation

August 14, 2026
저자: Yijiang Li, Yijun Liang, Yunjie Tian, Bingyang Wang, Ke Zhang, Zhenfei Yin, Di Fu, Philip Torr, Nuno Vasconcelos
cs.AI

초록

시각적 온-정책 증류는 더 크고 강한 교사 모델이나 참조 답안, 실측 관심 영역과 같은 특권적 감독을 통한 정보성 있는 교사-학생 비대칭성에 크게 의존한다. 이는 근본적인 질문을 제기한다: 특권적 정보가 전혀 없을 때 정보성 있는 비대칭성은 어디에서 비롯될 수 있는가? 우리는 비대칭성의 원천을 역전시켜 이 질문에 답한다. 교사에게 특권적 정보를 추가하는 대신, 학생으로부터 정보를 제거한다. 이러한 비대칭성은 실측 주석, 보상, 또는 별도의 더 강한 교사 모델 없이도, 학생이 접근할 수 없는 정보에 접근하는 교사와 동일한 효과적 학습 신호를 추가 비용 없이 생성한다. 이 원리를 바탕으로, 우리는 비대칭적 증강 뷰로부터 온-정책 학습 신호를 구성하는 간단하면서도 효과적인 방법인 자기지도 시각적 온-정책 증류(S^2VOPD)를 제안한다. S^2VOPD는 원본 이미지에 조건화된 교사 분포를 동일한 이미지의 강하게 증강된 뷰에 조건화된 학생 분포로 온-정책 증류한다. 우리는 시각적 증강의 광범위한 설계 공간을 체계적으로 탐구하여 다음을 발견한다: (1) 비대칭성이 중요하다: 네 가지 증강 패밀리 모두 성능을 향상시키는 반면, 대칭적 자기 증류는 성능을 저하시킨다; (2) 강도가 중요하다: 성능은 적절한 강도에서 최고조에 달한다; (3) 격차는 과제와 일관성을 유지해야 한다: 질문 관련 증거를 완전히 제거하는 증강은 크지만 정보를 제공하지 않는 불일치를 유발할 수 있다. 여섯 개의 세밀한 지각 벤치마크에서 S^2VOPD는 Qwen3.5-4B를 70.7%에서 77.4%로 향상시키며, 비교된 모든 오픈소스 모델(Qwen3-VL 235B까지 포함)보다 높은 성능을 달성하고 GPT-5.4를 능가한다. 훈련 데이터를 동일하게 유지하면서, 특권적 정보를 사용하는 방법들이 달성한 성능 향상의 96%를 회복한다. 웹사이트: https://williamium3000.github.io/s2vopd
English
Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest. This raises a fundamental question: where can informative asymmetry come from when nothing privileged is available? We answer this by inverting where the asymmetry comes from. Rather than adding privileged information to the teacher, we subtract information from the student. This asymmetry creates the same effective learning signal for free as a teacher with access to information unavailable to the student, without ground-truth annotations, rewards, or a separate stronger teacher model. Building on this principle, we introduce Self-Supervised Visual On-Policy Distillation (S^2VOPD), a simple yet effective method that constructs on-policy learning signals from asymmetric augmented views. S^2VOPD distills the teacher's distribution conditioned on the original image on-policy into the student distribution conditioned on a strongly augmented view of the same image. We systematically explore a broad design space of visual augmentations and uncover that (1) asymmetry matters: all four augmentation families improve performance, while symmetric self-distillation degrades it; (2) strength matters: performance peaks at a moderate strength; and (3) the gap must remain task-consistent: augmentations that completely remove the question-relevant evidence can induce large but uninformative discrepancies. Across six fine-grained perception benchmarks, S^2VOPD improves Qwen3.5-4B from 70.7% to 77.4%, above all open-source models compared, up to Qwen3-VL at 235B, and surpasses GPT-5.4. While holding training data the same, it recovers 96% of the improvement achieved by methods with privileged information. Website is at https://williamium3000.github.io/s2vopd