SIGNPOST-Bench: マルチモーダル大規模言語モデルにおけるテキストと視覚情報の競合解決のベンチマーキング
SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models
August 4, 2026
著者: Sirun Li, Minghao Liu, Ling Dai, Yong Li, Haoxin Lyu, Junting Zhou, Fan Zhang
cs.AI
要旨
マルチモーダル大規模言語モデル(MLLM)は、視覚的およびテキスト的手がかりを組み合わせることで、実世界のシーンにおける根拠に基づく予測を行う。しかし、既存のベンチマークでは、これらの証拠源が衝突した場合にモデルがどのように調停するかを明らかにすることはほとんどない。本稿では、テキストと視覚の競合解決を評価するための制御された反事実ベンチマークであるSIGNPOST-Benchを紹介する。各ソース画像は、オリジナル、ブランク、類似、ランダム、敵対的の五つのバリアントからなる反事実五つ組に変換される。非テキストコンテンツを保持するように設計された合成的で局所的なシーンテキスト介入により、位置特定性能の変化と、競合するテキストによって導入される地理的ターゲットへの方向付けられたシフトのペア測定が可能になる。SIGNPOST-Benchは、4つのデータセットから5,111の反事実グループと25,555の画像バリアントを含む。我々は、7つのプロバイダーから20のMLLMを評価した。オリジナル画像と比較して、敵対的バリアントは、位置特定誤差の中央値を282 kmから1,347 kmへと4.8倍増加させる。ジオコーディング可能な敵対的サンプルのうち、モデル全体で予測の6.5〜20.1%が注入されたターゲットから50 km未満に位置し、評価されたすべてのモデルにおいて、ブランクから敵対的へのターゲット距離の平均ペア削減が正の値を示した。適合するテキスト置換、無関係なテキスト置換、競合するテキスト置換は、モデルの予測に異なる影響を及ぼし、クリーンな入力での位置特定性能は、競合するテキストに対する頑健性を完全には予測しない。これらの結果は、視覚的地理位置特定がシーンテキスト調停の連続的な診断手段として機能することを実証し、MLLMが競合するマルチモーダル証拠をどのように解決するかを評価するための制御された枠組みを提供する。
English
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.