ChatPaper.aiChatPaper

MMOOC: マルチモーダル大規模言語モデルにおける文脈外評価のための包括的ベンチマーク

MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

August 1, 2026
著者: Wenjie Zhu, Yabin Zhang, Wenjun Zeng, Lei Zhang
cs.AI

要旨

マルチモーダル大規模言語モデル(MLLMs)は、多様な視覚言語タスクにおいて高い性能を達成しているが、不完全またはシフトした文脈の下ではしばしば失敗する。信頼性の高いMLLMは、主題レベルの文脈シフトを伴う真に文脈外(OOC)の質問には回答を拒否すべきであり、一方で非主題の文脈シフトを伴うシフトした文脈内(Shifted IC)の質問には引き続き回答すべきである。既存のベンチマークは主にOOCまたは視覚的に回答不可能な質問に焦点を当てているが、回答可能なShifted ICケースを見落としており、扱っているOOCシフトも限定的である。このギャップを埋めるため、我々はMLLMの拒否能力と頑健な回答能力を評価するための大規模ベンチマークであるMMOOCを提案する。MMOOCは、回答可能なShifted ICケースと回答不可能なOOCケースを含む41,000以上の画像・質問ペアから構成され、3つの質問形式、8つのシフトタイプ、6つの視覚シナリオにわたり、データ品質はMLLMベースのフィルタリングと人間による検証を通じて確保されている。我々は、AccuracyとRefusal Rateを用いてモデルの応答を評価し、さらにモデルの推論の正しさを評価するためのLLM-as-a-Judgeメトリクスを導入する。多様なMLLMに対する実験により、現在のモデルはシフトした文脈の下で回答可能性と拒否のバランスを取ることに依然として苦戦していることが示される。さらに、主要な失敗パターンを分析し、ポストトレーニングが頑健性を向上させ得ることを示す。MMOOCは公開される予定である。
English
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook answerable Shifted IC cases and cover limited OOC shifts. To fill this gap, we present MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs. MMOOC contains over 41K image-question pairs, including answerable Shifted IC cases and unanswerable OOC cases, spanning three question formats, eight shift types and six visual scenarios, with data quality ensured through MLLM-based filtering and human verification. We evaluate model responses using Accuracy and Refusal Rate, and further introduce an LLM-as-a-Judge metric to assess the correctness of model reasoning. Experiments on diverse MLLMs show that current models still struggle to balance answer-ability and refusal under shifted contexts. We further analyze key failure patterns and show that post-training can improve robustness. MMOOC will be made publicly available.