ChatPaper.aiChatPaper

SIGNPOST-Bench:多模态大语言模型中文本-视觉冲突消解的基准评测

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

August 4, 2026
作者: Sirun Li, Minghao Liu, Ling Dai, Yong Li, Haoxin Lyu, Junting Zhou, Fan Zhang
cs.AI

摘要

多模态大语言模型(MLLMs)通过结合视觉和文本线索在真实场景中进行有根据的预测,然而现有基准很少揭示当这些证据来源发生冲突时模型如何在它们之间进行仲裁。我们提出SIGNPOST-Bench,一个用于评估文本-视觉冲突消解的可控反事实基准。每张源图像被转换为由原始(Original)、空白(Blank)、相似(Similar)、随机(Random)和对抗(Adversarial)变体组成的反事实五元组。合成的局部场景文本干预旨在保留非文本内容,从而能够对定位性能的变化以及由冲突文本引入的针对地理目标的定向偏移进行配对测量。SIGNPOST-Bench包含来自四个数据集的5,111个反事实组和25,555个图像变体。我们评估了来自七个提供商的20个MLLM。与原始图像相比,对抗变体将中位定位误差从282公里提高到1,347公里,增加了4.8倍。在可地理编码的对抗样本中,各模型有6.5%-20.1%的预测距离注入目标不到50公里,并且每个被评估模型都表现出从空白到对抗变体在目标距离上的正向平均配对缩减。兼容、无关和冲突的文本替换对模型预测产生不同的效应,而干净输入下的定位性能并不能完全预测对冲突文本的鲁棒性。这些结果将视觉地理定位确立为场景文本仲裁的连续诊断工具,并为评估MLLM如何解决冲突的多模态证据提供了一个受控框架。
English
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.