VAD:多模态在策略蒸馏中面向目标重建的视觉证据归因
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
July 30, 2026
作者: Kangning Zhang, Yixing Li, Shuai Shao, Qingyao Li, Zhengxi Lu, Zhiyuan Yao, Jianghao Lin, Wenxiang Jiao, Yuan Lu, Weiwen Liu, Weinan Zhang, Yong Yu
cs.AI
摘要
多模态在线策略蒸馏(OPD)通过特权视角教师对学生生成轨迹进行监督,从而迁移细粒度的视觉知识。然而,其下一词元修正属于源自混合性质,将视觉信号与语言先验及教师特定效应混杂在一起。关键挑战在于估计哪些修正受到视觉证据支持,而不仅仅是确定在何处或以多大强度进行蒸馏。我们提出视觉归因蒸馏(VAD),一种反事实目标重建算法,用于估计教师修正中可由视觉归因的部分。在每个学生生成的前缀处,VAD在相关证据存在与移除两种情形下评估同一固定教师。中心化对数概率的相应变化定义了u_t,即视觉证据方向的带符号代理,用于估计证据对候选词元是支持还是反驳的揭示程度。VAD将原始修正投影到该代理上,获得干预对齐分量和代理未解释残差,然后从前者的信息重建出学生锚定目标。在训练过程中,该重建目标提供主要监督信号,而特权视角教师仅贡献弱正则化。在4B和9B规模的六个细粒度视觉基准上,VAD优于直接特权视角蒸馏和视觉优势加权。词元级与受控目标分析表明,代理对齐分量富含任务相关的视觉修正,并能产生更强的目标偏移,尤其在证据反驳错误答案时效果更为显著。这些结果支持反事实目标重建作为源自混合监督的有效替代方案。
English
Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.