ChatPaper.aiChatPaper

OPD-V:具有模态平衡的视觉在线策略自蒸馏

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

August 5, 2026
作者: Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
cs.AI

摘要

同策略自蒸馏(OPSD)已成为提升多模态大语言模型(MLLMs)视觉推理能力的标准后训练方法。现有方法从多种输入来源中提取特权信息来引导自蒸馏,但这些设计忽略了模态不平衡——一个MLLM推理中固有的挑战。当文本信息主导生成过程时,模型无法充分整合其多模态输入,因而精心设计的特权信息仍未得到充分利用,限制了OPSD的有效性。为探究这一局限,我们构建了使用放大图像的正教师和使用掩码图像的负教师,二者表现出不同程度的模态不平衡。它们的推理正确性与token logits的变化揭示出:模态平衡本身即可作为特权信息。受此发现启发,我们提出OPD-V,一种通过正教师和负教师实例化该信息的视觉OPSD范式。正教师的模态平衡Logits边际定义了一个模态平衡信任区域,用于选择进行自蒸馏的同策略词元。在6个基准、4个MLLM骨干网络和5种后训练方法上的实验表明,OPD-V在降低训练成本的同时持续提升了推理性能。
English
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-distillation. Yet these designs overlook Modality Imbalance, a challenge inherent to MLLM reasoning. When textual information dominates generation, the model cannot fully integrate its multimodal input. Consequently, carefully designed privileged information remains underused, limiting the effectiveness of OPSD. To examine this limitation, we construct a Positive Teacher with the Zoom-In Image and a Negative Teacher with the Mask Image, which exhibit different degrees of Modality Imbalance. Changes in their reasoning correctness and token logits reveal that Modality Balance can itself serve as privileged information. Motivated by this finding, we introduce OPD-V, a visual OPSD paradigm that instantiates such information through the Positive Teacher and Negative Teacher. Positive Modality-Balance Logits Margins define a Modality-Balance Trust Region that selects the on-policy tokens used for self-distillation. Experiments across 6 benchmarks, 4 MLLM backbones, and 5 post-training methods show that OPD-V consistently improves reasoning performance while reducing training cost.