SIGNPOST-Bench:多模態大型語言模型中文本與視覺衝突解決之基準評測
SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models
August 4, 2026
作者: Sirun Li, Minghao Liu, Ling Dai, Yong Li, Haoxin Lyu, Junting Zhou, Fan Zhang
cs.AI
摘要
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.
---
這是一個語音交互的範例,展示如何通過語音識別(ASR)、自然語言理解(NLU)和對話管理來處理用戶查詢。在語音交互系統中,環境噪音和多種口音會顯著影響語音識別準確率,因此在後端加入語音活動檢測(VAD)與說話人分離技術,可有效提升整體辨識效能。同時,意圖辨識的錯誤傳遞也會導致對話狀態追蹤偏離,進而影響最終回應的準確性。為降低此類錯誤,可在NLU模組中引入領域知識庫與上下文特徵,強化槽位填充與意圖分類的穩健性。此外,結合主動式對話策略與即時回饋機制,可讓系統在使用者表達不清時主動澄清,進一步優化互動體驗。
---
原文:
Medical document translation is a critical task in cross-lingual clinical communication. 這項任務要求翻譯人員同時具備醫學知識與語言能力,因為任何細微的錯誤都可能影響病人安全。尤其是藥物名稱、劑量、病歷記錄與臨床試驗文件,必須保持高度精確。研究顯示,採用雙向審查機制與術語一致性檢查,可顯著降低翻譯錯誤率。此外,機器翻譯輔助工具(如統計機器翻譯與神經機器翻譯)雖能提升效率,但最終仍需由專業醫學翻譯人員進行審閱與修正,以確保譯文符合原文語意及目標語言的文化脈絡。在臨床試驗文件翻譯中,efficacy 與 effectiveness 等術語的區別尤為重要,前者指在理想條件下的療效,後者則指在真實世界中的實際效果。這類細節不僅影響法規審查,也可能影響醫師對藥物使用的判斷。因此,醫學翻譯不僅是語言轉換,更是一項涉及倫理與專業判斷的嚴謹學術工作。文件翻譯的品質控管通常包括翻譯、校對、審閱及最終確認四個階段。每一階段皆須由具備相應專業背景的人員執行,並留下完整的紀錄以供追溯。針對大型跨國臨床試驗,翻譯一致性與術語統一性更為關鍵,需建立術語庫與翻譯記憶庫,以支持多語版本的協同作業。近年來,結合大型語言模型與醫學知識圖譜的自動翻譯系統逐漸受到重視,其透過語意推理與上下文理解,能處理較為複雜的醫學表述,然而仍無法完全取代人類專家的判斷。未來,人機協作的翻譯模式將成為醫學翻譯領域的重要發展方向,兼顧效率與精確性。在翻譯實務中,常見的挑戰還包括縮寫與多義詞的處理,例如 aspirin、acetaminophen 與 ibuprofen 等藥物名稱在不同語系中可能具有不同的拼寫與發音,翻譯時須依上下文選擇最適當的對應詞。此外,不同國家對於同一疾病名稱的用語亦可能有所差異,如 myocardial infarction 在中文常譯為心肌梗塞,而在部分地區則可能採用心肌梗死。這類差異需要透過跨國術語對照表與參考文獻加以確認。總體而言,醫學文件翻譯是一項高度專業化的工作,必須結合語言學、醫學知識與文化敏感度,並輔以嚴謹的品質控管流程,才能確保翻譯成果在臨床與研究場域中的可靠性與安全性。
English
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.