VAD:多模態在策略蒸餾中目標重建的視覺證據歸因
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
July 30, 2026
作者: Kangning Zhang, Yixing Li, Shuai Shao, Qingyao Li, Zhengxi Lu, Zhiyuan Yao, Jianghao Lin, Wenxiang Jiao, Yuan Lu, Weiwen Liu, Weinan Zhang, Yong Yu
cs.AI
摘要
多模態在策略蒸餾(OPD)透過特權視圖教師監督學生生成的軌跡,來遷移細粒度的視覺知識。然而,其下一個詞元的修正是來源混合的,結合了視覺信號、語言先驗與教師特定效應。關鍵挑戰在於估計哪些修正受到視覺證據的支持,而不僅是決定蒸餾的位置或強度。我們提出視覺歸因蒸餾(VAD),這是一種反事實目標重建演算法,用於估計教師修正中視覺可歸因的部分。在每個學生生成的前綴處,VAD 在相關證據存在與移除兩種情況下評估同一個固定教師。對應的中心化對數機率變化定義了 u_t,這是視覺證據方向的帶符號代理,用以估計證據支持或反駁候選詞元的揭示程度。VAD 將原始修正投影到此代理上,獲得干預對齊分量與代理未解釋殘差,再從前者重建出學生錨定的目標。在訓練期間,此重建目標提供主要的監督信號,而特權教師則貢獻一個弱正則化器。在 4B 與 9B 規模的六個細粒度視覺基準上,VAD 優於直接特權視圖蒸餾與視覺優勢加權。詞元層級與受控目標分析顯示,代理對齊分量富含與任務相關的視覺修正,並產生更強的目標偏移,尤其是當證據反駁錯誤答案時。這些結果支持反事實目標重建作為來源混合監督的有效替代方案。
English
Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.