ChatPaper.aiChatPaper

보는 것인가, 아는 것인가? 멀티모달 대규모 언어 모델의 시각적 맥락 민감성

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

July 28, 2026
저자: Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, Vésteinn Snæbjarnarson
cs.AI

초록

다중모달 대규모 언어 모델(MLLM)은 사전 학습된 언어 모델의 풍부한 사전 지식과 시각적 입력을 통합함으로써 강력한 성능을 달성한다. 그러나 이러한 모델은 시각적 증거가 사전 학습된 지식과 충돌할 때, 특히 시각 중심 작업에서 종종 실패한다. 우리는 두 가지 진단 패러다임을 사용하여 이러한 실패를 개별적으로 탐구한다. (1) 이미지 재구성을 통해 시각적 정보가 이용 가능한지 조사하고, (2) 모델이 언어 사전보다 시각적 맥락을 따르는 정도인 다중모달 맥락 민감도를 측정한다. 두 번째를 지원하기 위해 우리는 다섯 가지 거친 수준의 차원(공간-시간, 색상, 개수, 크기, 무게)에 걸친 WhatIfVis 벤치마크를 도입하며, 이 벤치마크의 질문은 이미지 또는 사전 지식 중 어느 쪽에서도 답을 얻을 수 있다. 우리의 분석은 세 가지 발견을 제시한다. (i) 거친 수준의 시각적 증거는 보존되는데, 이는 이러한 속성들이 고정된(frozen) MLLM의 최종 계층 이미지 토큰에서 재구성될 수 있기 때문이다. 따라서 이러한 속성에 관한 질문에서의 실패는 지각 중 시각적 인코딩의 저하가 아니라 지각 후 활용 문제를 가리킨다. (ii) 시각적 증거를 사용하거나 무시하도록 명시적으로 지시받은 경우에도 (WhatIfVis에 대한 지도 미세 조정을 거치지 않은) 기본(vanilla) 모델은 불안정한 시각적 맥락 민감도를 보인다. 지도 미세 조정(SFT)은 이러한 제어 가능성을 개선하고 도메인 전반에 걸쳐 일반화되며, 활성화 패칭(activation patching)은 6개 모델 모두에서 시각-대-사전 지식 간의 절충이 발생하는 아키텍처별 깊이를 더욱 정밀하게 특정한다. (iii) 시각-대-사전 지식 간의 절충은 학습된 벡터를 따라 제어할 수 있다. 이 조향 벡터(steering vector)를 의도 명령 없이 적용하더라도 기본 모델에 비해 제어 가능성이 향상된다. 종합적으로, 이러한 결과는 병목 지점을 재규정하며, 이는 우리가 연구한 거친 속성에 대해 MLLM이 시각적 증거를 인코딩하지만 이에 대한 의존성을 안정적으로 제어할 수 없음을 나타낸다.
English
Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visual evidence conflicts with pretrained knowledge. We explore these failures separately using two diagnostic paradigms: (1) probing whether visual information is available, via image reconstruction, and (2) measuring multimodal context sensitivity, the extent to which the model follows visual context versus the language prior. To support the second, we introduce the WhatIfVis, a benchmark spanning five coarse-grained dimensions (spatial-temporal, color, count, size, and weight) whose questions admit answers from either the image or the prior. Our analysis yields three findings: (i) Coarse-grained visual evidence is preserved, as these attributes can be reconstructed from the final-layer image tokens of frozen MLLMs. Failures on questions about these attributes therefore point to post-perceptual utilization, rather than to degraded visual encoding during perception. (ii) Even when explicitly instructed to use or ignore visual evidence, vanilla models (without supervised fine-tuning on the WhatIfVis) show unstable visual context sensitivity. Supervised fine-tuning (SFT) improves this controllability and generalizes across domains, and activation patching further localizes the vision-versus-prior trade-off at architecture-specific depths across all six models. (iii) The vision-versus-prior trade-off is controllable along a learned vector. Applying this steering vector, even without any intent instruction, improves controllability over the vanilla model. Together, these results relocate the bottleneck, indicating that for the coarse attributes we study, MLLMs encode the visual evidence but cannot reliably control their reliance on it.