ChatPaper.aiChatPaper

OPD-V:視覺在線策略自蒸餾與模態平衡

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

August 5, 2026
作者: Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
cs.AI

摘要

同策略自蒸餾(On-Policy Self-Distillation, OPSD)已成為提升多模態大型語言模型(MLLMs)視覺推理能力的標準後訓練方法。現有方法從多種輸入來源提取特權信息以引導自蒸餾。然而,這些設計忽略了模態不平衡——一個MLLM推理中固有的挑戰。當文本信息主導生成時,模型無法充分整合其多模態輸入。因此,精心設計的特權信息仍未被充分利用,限制了OPSD的有效性。為檢視此限制,我們以放大圖像建構正向教師,並以遮罩圖像建構負向教師,兩者展現不同程度的模態不平衡。其推理正確性與詞元邏輯值的變化顯示,模態平衡本身即可作為特權信息。受此發現啟發,我們提出OPD-V,一種視覺OPSD範式,透過正向教師與負向教師具體化此類信息。正向模態平衡邏輯值邊際定義了模態平衡信任區域,用以選取用於自蒸餾的同策略詞元。在6個基準、4個MLLM骨幹模型與5種後訓練方法上的實驗顯示,OPD-V在降低訓練成本的同時,持續提升推理表現。
English
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-distillation. Yet these designs overlook Modality Imbalance, a challenge inherent to MLLM reasoning. When textual information dominates generation, the model cannot fully integrate its multimodal input. Consequently, carefully designed privileged information remains underused, limiting the effectiveness of OPSD. To examine this limitation, we construct a Positive Teacher with the Zoom-In Image and a Negative Teacher with the Mask Image, which exhibit different degrees of Modality Imbalance. Changes in their reasoning correctness and token logits reveal that Modality Balance can itself serve as privileged information. Motivated by this finding, we introduce OPD-V, a visual OPSD paradigm that instantiates such information through the Positive Teacher and Negative Teacher. Positive Modality-Balance Logits Margins define a Modality-Balance Trust Region that selects the on-policy tokens used for self-distillation. Experiments across 6 benchmarks, 4 MLLM backbones, and 5 post-training methods show that OPD-V consistently improves reasoning performance while reducing training cost.