MMOOC:多模態大型語言模型中上下文外評估的綜合基準
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
August 1, 2026
作者: Wenjie Zhu, Yabin Zhang, Wenjun Zeng, Lei Zhang
cs.AI
摘要
多模態大型語言模型(MLLMs)在廣泛的視覺語言任務上已展現出強勁的表現,但在不完美或情境偏移的條件下往往會失效。一個可靠的 MLLM 應能拒答真正上下文不符(OOC)且涉及主題層級情境偏移的問題,同時仍能回答帶有非主題情境偏移的偏移但仍在上下文內(Shifted IC)問題。現有基準測試主要針對 OOC 或視覺上無法回答的問題,卻忽略了可回答的 Shifted IC 案例,且涵蓋的 OOC 偏移類型有限。為填補此缺口,我們提出 MMOOC,一個大規模評估 MLLMs 拒答能力與穩健回答能力的基準測試。MMOOC 包含超過 41K 組影像-問題配對,涵蓋可回答的 Shifted IC 案例與不可回答的 OOC 案例,橫跨三種問題格式、八種偏移類型及六種視覺場景,並透過基於 MLLM 的篩選與人工驗證確保資料品質。我們使用準確率(Accuracy)與拒答率(Refusal Rate)評估模型回應,並進一步引入 LLM-as-a-Judge 指標來評估模型推理的正確性。在多家不同 MLLMs 上的實驗顯示,當前模型在情境偏移下仍難以在可回答性與拒答之間取得平衡。我們進一步分析了主要的失敗模式,並顯示後訓練(post-training)能提升穩健性。MMOOC 將公開釋出。
English
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook answerable Shifted IC cases and cover limited OOC shifts. To fill this gap, we present MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs. MMOOC contains over 41K image-question pairs, including answerable Shifted IC cases and unanswerable OOC cases, spanning three question formats, eight shift types and six visual scenarios, with data quality ensured through MLLM-based filtering and human verification. We evaluate model responses using Accuracy and Refusal Rate, and further introduce an LLM-as-a-Judge metric to assess the correctness of model reasoning. Experiments on diverse MLLMs show that current models still struggle to balance answer-ability and refusal under shifted contexts. We further analyze key failure patterns and show that post-training can improve robustness. MMOOC will be made publicly available.