OPD-V: モダリティバランスを備えた視覚的オン・ポリシー自己蒸留
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
August 5, 2026
著者: Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
cs.AI
要旨
オン方策自己蒸留(OPSD)は、マルチモーダル大規模言語モデル(MLLM)における視覚的推論を向上させるための標準的なポストトレーニング手法となっている。既存手法は、多様な入力ソースから得られる特権情報を引き出し、自己蒸留を導いている。しかしながら、これらの設計は、MLLM推論に内在する課題であるモダリティ不均衡を見落としている。テキスト情報が生成を支配する場合、モデルはマルチモーダル入力を十分に統合できない。その結果、慎重に設計された特権情報は十分に活用されず、OPSDの有効性が制限される。この限界を検証するため、我々はズームイン画像を用いたポジティブ教師とマスク画像を用いたネガティブ教師を構築し、これらが異なる程度のモダリティ不均衡を示すことを確認した。それらの推論正確性とトークンロジットの変化から、モダリティバランス自体が特権情報として機能し得ることが明らかになった。この知見に基づき、我々はポジティブ教師とネガティブ教師を通じてそのような情報を具体化する視覚的OPSDパラダイムであるOPD-Vを導入する。ポジティブなモダリティバランスのロジットマージンは、自己蒸留に使用するオン方策トークンを選択するモダリティバランス信頼領域を定義する。6つのベンチマーク、4つのMLLMバックボーン、5つのポストトレーニング手法にわたる実験により、OPD-Vが学習コストを削減しながら推論性能を一貫して向上させることが示された。
English
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-distillation. Yet these designs overlook Modality Imbalance, a challenge inherent to MLLM reasoning. When textual information dominates generation, the model cannot fully integrate its multimodal input. Consequently, carefully designed privileged information remains underused, limiting the effectiveness of OPSD. To examine this limitation, we construct a Positive Teacher with the Zoom-In Image and a Negative Teacher with the Mask Image, which exhibit different degrees of Modality Imbalance. Changes in their reasoning correctness and token logits reveal that Modality Balance can itself serve as privileged information. Motivated by this finding, we introduce OPD-V, a visual OPSD paradigm that instantiates such information through the Positive Teacher and Negative Teacher. Positive Modality-Balance Logits Margins define a Modality-Balance Trust Region that selects the on-policy tokens used for self-distillation. Experiments across 6 benchmarks, 4 MLLM backbones, and 5 post-training methods show that OPD-V consistently improves reasoning performance while reducing training cost.