ChatPaper.aiChatPaper

VAD: マルチモーダル・オンポリシー蒸留におけるターゲット再構成のための視覚的証拠の帰属

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

July 30, 2026
著者: Kangning Zhang, Yixing Li, Shuai Shao, Qingyao Li, Zhengxi Lu, Zhiyuan Yao, Jianghao Lin, Wenxiang Jiao, Yuan Lu, Weiwen Liu, Weinan Zhang, Yong Yu
cs.AI

要旨

マルチモーダル・オンポリシー蒸留(OPD)は、特権視点を持つ教師モデルを用いて学生が生成した軌跡を監督することで、細粒度の視覚知識を転移する。しかし、その次トークン修正はソース混合であり、視覚信号と言語的な事前知識、さらには教師固有の効果が組み合わさっている。重要な課題は、どの修正が視覚的証拠によって裏付けられているかを推定することであり、単にどこで、どの程度の強さで蒸留するかではない。本稿では、教師修正のうち視覚に帰属可能な部分を推定する反事実的ターゲット再構成アルゴリズムであるVisual Attribution Distillation(VAD)を提案する。各学生生成プレフィックスにおいて、VADは関連する証拠が存在する場合と取り除かれた場合の同一の固定教師を評価する。対応する中心化ログ確率の変化はu_tを定義する。これは視覚的証拠の方向を示す符号付きプロキシであり、証拠が候補トークンをどの程度支持または反駁するかを推定する。VADは元の修正をこのプロキシに投影し、介入整合成分とプロキシ未説明残差を取得し、前者から学生に固定されたターゲットを再構成する。訓練中、この再構成ターゲットが主要な監督信号を提供し、特権教師は弱い正則化項として寄与する。4Bおよび9B規模の6つの細粒度視覚ベンチマークにおいて、VADは直接的な特権視点蒸留や視覚的優位性重み付けよりも優れた性能を示した。トークンレベルおよび制御ターゲットの分析により、プロキシ整合成分はタスク関連の視覚修正に富み、特に証拠が誤った回答を反駁する場合に、より強いターゲットシフトをもたらすことが明らかになった。これらの結果は、反事実的ターゲット再構成がソース混合監督の効果的な代替手段であることを支持するものである。
English
Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.