VAD: 멀티모달 온-폴리시 증류에서 대상 재구성을 위한 시각적 증거 귀속
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
July 30, 2026
저자: Kangning Zhang, Yixing Li, Shuai Shao, Qingyao Li, Zhengxi Lu, Zhiyuan Yao, Jianghao Lin, Wenxiang Jiao, Yuan Lu, Weiwen Liu, Weinan Zhang, Yong Yu
cs.AI
초록
멀티모달 온-폴리시 증류(OPD)는 특권 시점 교사(privileged-view teacher)로 학생 생성 궤적을 지도하여 세밀한 시각 지식을 전이한다. 그러나 그 다음 토큰 교정은 소스 혼합(source-mixed) 방식으로, 시각 신호에 언어적 사전 지식과 교사 특유의 효과가 결합되어 있다. 핵심 과제는 어디서 또는 얼마나 강하게 증류할지가 아니라, 어떤 교정이 시각적 증거에 의해 뒷받침되는지를 추정하는 것이다. 우리는 교정의 시각적 귀인이 가능한 부분을 추정하는 반사실적 대상 재구성 알고리즘인 시각적 속성 증류(VAD)를 도입한다. 각 학생 생성 접두사에서 VAD는 관련 증거를 포함한 상태와 제거한 상태에서 동일한 고정 교사를 평가한다. 중심화된 로그 확률의 해당 변화는 ut를 정의하며, 이는 증거가 후보 토큰을 지지하거나 반박하는 정도를 추정하는 시각적 증거 방향의 부호화된 프록시 역할을 한다. VAD는 원래 교정을 이 프록시에 투영하여 개입 정렬 성분과 프록시로 설명되지 않는 잔차를 얻은 후, 전자에서 학생 중심 대상을 재구성한다. 훈련 중 이 재구성된 대상이 주요 지도 신호를 제공하고, 특권 교사는 약한 정규화기로 기여한다. 4B 및 9B 규모의 여섯 가지 세밀한 시각 벤치마크에서 VAD는 직접 특권 시점 증류 및 시각적 이점 가중치보다 우수한 성능을 보인다. 토큰 수준 및 통제 대상 분석은 프록시 정렬 성분이 과제 관련 시각 교정에 풍부하며, 특히 증거가 오답을 반박할 때 더 강한 대상 이동을 생성함을 보여준다. 이러한 결과는 반사실적 대상 재구성이 소스 혼합 지도의 효과적인 대안임을 뒷받침한다.
English
Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.