SIGNPOST-Bench: 다중모달 대규모 언어 모델에서의 텍스트-비전 갈등 해결 능력 벤치마킹
SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models
August 4, 2026
저자: Sirun Li, Minghao Liu, Ling Dai, Yong Li, Haoxin Lyu, Junting Zhou, Fan Zhang
cs.AI
초록
다중 모드 대규모 언어 모델(MLLM)은 시각적 단서와 텍스트적 단서를 결합하여 실제 세계 장면에 근거한 예측을 수행하지만, 기존 벤치마크는 이러한 증거 소스들이 충돌할 때 모델이 어떻게 중재하는지를 거의 드러내지 않는다. 우리는 텍스트-시각 충돌 해결을 평가하기 위한 통제된 반사실적 벤치마크인 SIGNPOST-Bench를 소개한다. 각 원본 이미지는 원본(Original), 공백(Blank), 유사(Similar), 무작위(Random), 적대적(Adversarial) 변형으로 구성된 반사실적 다섯 가지 변형 세트로 변환된다. 비텍스트 콘텐츠를 보존하도록 설계된 합성적이고 국소화된 장면-텍스트 개입을 통해, 위치 파악 성능의 변화와 충돌하는 텍스트가 유도하는 지리적 목표를 향한 방향적 이동을 쌍별로 측정할 수 있다. SIGNPOST-Bench는 네 개의 데이터셋에서 추출한 5,111개의 반사실적 그룹과 25,555개의 이미지 변형을 포함한다. 우리는 일곱 개 제공업체의 20개 MLLM을 평가한다. 원본 이미지와 비교할 때, 적대적 변형은 위치 파악 오차의 중앙값을 282km에서 1,347km로, 즉 4.8배 증가시킨다. 지오코딩 가능한 적대적 샘플 중에서, 모델에 따라 예측의 6.5~20.1%가 주입된 목표로부터 50km 이내에 위치한다. 또한 평가된 모든 모델은 Blank에서 Adversarial로 전환될 때 목표 거리에서 양의 평균 쌍별 감소를 보인다. 일치하는, 무관한, 그리고 충돌하는 텍스트 대체는 모델 예측에 뚜렷하게 다른 영향을 미친다. 반면, 정상 입력에서의 위치 파악 성능은 충돌하는 텍스트에 대한 견고성을 완전히 예측하지는 못한다. 이러한 결과는 시각적 지리 위치 파악을 장면-텍스트 중재의 연속적 진단 도구로 확립하며, MLLM이 충돌하는 다중 모드 증거를 어떻게 해결하는지 평가하기 위한 통제된 프레임워크를 제공한다.
English
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.