OPD-V: 모달리티 균형을 통한 시각적 온-폴리시 자기 증류
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
August 5, 2026
저자: Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
cs.AI
초록
온-폴리시 자기 증류(OPSD)는 다중모달 대규모 언어 모델(MLLM)의 시각적 추론을 개선하기 위한 표준적인 사후 훈련 접근법이 되었다. 기존 방법들은 자기 증류를 안내하기 위해 다양한 입력 소스에서 특권 정보를 끌어낸다. 그러나 이러한 설계는 MLLM 추론에 내재된 도전 과제인 모달리티 불균형(Modality Imbalance)을 간과한다. 텍스트 정보가 생성을 지배할 때 모델은 다중모달 입력을 완전히 통합하지 못한다. 결과적으로 신중하게 설계된 특권 정보는 충분히 활용되지 못하며, 이는 OPSD의 효과성을 제한한다. 이 한계를 조사하기 위해 우리는 확대 이미지(Zoom-In Image)를 사용한 긍정 교사(Positive Teacher)와 마스크 이미지(Mask Image)를 사용한 부정 교사(Negative Teacher)를 구축한다. 이들은 서로 다른 정도의 모달리티 불균형을 나타낸다. 이들의 추론 정확도와 토큰 로짓의 변화는 모달리티 균형 자체가 특권 정보로 작용할 수 있음을 보여준다. 이 발견에 착안하여, 우리는 긍정 교사와 부정 교사를 통해 그러한 정보를 구체화하는 시각적 OPSD 패러다임인 OPD-V를 소개한다. 긍정 모달리티 균형 로짓 마진(Positive Modality-Balance Logits Margins)은 자기 증류에 사용되는 온-폴리시 토큰을 선택하는 모달리티 균형 신뢰 영역(Modality-Balance Trust Region)을 정의한다. 6개 벤치마크, 4개 MLLM 백본, 5개 사후 훈련 방법에 걸친 실험 결과, OPD-V는 훈련 비용을 줄이면서 추론 성능을 일관되게 향상시키는 것으로 나타났다.
English
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-distillation. Yet these designs overlook Modality Imbalance, a challenge inherent to MLLM reasoning. When textual information dominates generation, the model cannot fully integrate its multimodal input. Consequently, carefully designed privileged information remains underused, limiting the effectiveness of OPSD. To examine this limitation, we construct a Positive Teacher with the Zoom-In Image and a Negative Teacher with the Mask Image, which exhibit different degrees of Modality Imbalance. Changes in their reasoning correctness and token logits reveal that Modality Balance can itself serve as privileged information. Motivated by this finding, we introduce OPD-V, a visual OPSD paradigm that instantiates such information through the Positive Teacher and Negative Teacher. Positive Modality-Balance Logits Margins define a Modality-Balance Trust Region that selects the on-policy tokens used for self-distillation. Experiments across 6 benchmarks, 4 MLLM backbones, and 5 post-training methods show that OPD-V consistently improves reasoning performance while reducing training cost.