「見ること」か「知ること」か?——マルチモーダル大規模言語モデルにおける視覚的文脈感受性
Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
July 28, 2026
著者: Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, Vésteinn Snæbjarnarson
cs.AI
要旨
マルチモーダル大規模言語モデル(MLLM)は、視覚入力を事前学習済み言語モデルの豊富な事前知識と統合することで高い性能を達成する。しかし、視覚中心のタスク、特に視覚的証拠が事前知識と矛盾する場合には、しばしば失敗する。我々は、この失敗を2つの診断パラダイムを用いて個別に調査する。(1) 画像再構成による視覚情報の利用可能性の調査、(2) マルチモーダル文脈感度、すなわちモデルが言語事前知識よりも視覚的文脈にどの程度従うかの測定である。2つ目の調査を支援するために、我々はWhatIfVisを導入する。これは5つの粗粒度の次元(時空間、色、数、大きさ、重さ)にわたるベンチマークであり、その質問は画像または事前知識のいずれかから答えを得ることができる。
我々の分析は3つの知見をもたらす。(i) 粗粒度の視覚的証拠は保持されている。なぜなら、これらの属性は凍結されたMLLMの最終層の画像トークンから再構成できるからである。したがって、これらの属性に関する質問での失敗は、知覚中の視覚エンコーディングの劣化ではなく、知覚後の利用の問題を示している。(ii) 視覚的証拠を使用または無視するよう明示的に指示された場合でも、バニラモデル(WhatIfVisでの教師ありファインチューニングを行っていないモデル)は不安定な視覚的文脈感度を示す。教師ありファインチューニング(SFT)はこの制御可能性を向上させ、ドメイン間で一般化する。さらに、アクティベーションパッチングにより、全6モデルにおいて、視覚対事前知識のトレードオフがアーキテクチャ固有の深さに局在することが明らかになる。(iii) 視覚対事前知識のトレードオフは、学習されたベクトルに沿って制御可能である。このステアリングベクトルを適用すると、意図の指示がなくても、バニラモデルよりも制御可能性が向上する。これらの結果を総合すると、ボトルネックは移動する。すなわち、我々が研究する粗い属性に関して、MLLMは視覚的証拠をエンコードするが、それへの依存を確実に制御できないことを示している。
English
Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visual evidence conflicts with pretrained knowledge. We explore these failures separately using two diagnostic paradigms: (1) probing whether visual information is available, via image reconstruction, and (2) measuring multimodal context sensitivity, the extent to which the model follows visual context versus the language prior. To support the second, we introduce the WhatIfVis, a benchmark spanning five coarse-grained dimensions (spatial-temporal, color, count, size, and weight) whose questions admit answers from either the image or the prior. Our analysis yields three findings: (i) Coarse-grained visual evidence is preserved, as these attributes can be reconstructed from the final-layer image tokens of frozen MLLMs. Failures on questions about these attributes therefore point to post-perceptual utilization, rather than to degraded visual encoding during perception. (ii) Even when explicitly instructed to use or ignore visual evidence, vanilla models (without supervised fine-tuning on the WhatIfVis) show unstable visual context sensitivity. Supervised fine-tuning (SFT) improves this controllability and generalizes across domains, and activation patching further localizes the vision-versus-prior trade-off at architecture-specific depths across all six models. (iii) The vision-versus-prior trade-off is controllable along a learned vector. Applying this steering vector, even without any intent instruction, improves controllability over the vanilla model. Together, these results relocate the bottleneck, indicating that for the coarse attributes we study, MLLMs encode the visual evidence but cannot reliably control their reliance on it.