Blind-Spots-Bench: マルチモーダルモデルにおける盲点の評価
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
July 9, 2026
著者: Matteo Santelmo, Xiuying Wei, Israa Fakih, Felix Bauer, Juan Garcia Giraldo, Chengkun Li, Etienne Bamas, Emmanuel Abbé
cs.AI
要旨
最新のAIモデルは、多くの確立されたベンチマークで優れた性能を達成しているものの、文字列操作や5本足の犬を描くといった、人間にとってはほぼ自明と感じられるタスクでは依然として失敗する。これらの例は、既存のベンチマークが現在のシステムにおける持続的な盲点を過小評価している可能性を示唆している。本稿では、人間には単純に見えるが最新のAIには困難なタスクを通じて、そのような盲点を明らかにするために設計されたベンチマーク「blind-spots-bench」を導入する。我々は、AIコースの学生から収集した素の質問を整理し、構造化された参照解を付与して注釈を付け、結果として得られた235サンプルのデータセットに適合したタスク分類法を提案する。さらに、オープンウェイトおよびクローズドソースの言語モデル、視覚言語モデル、画像生成モデルを含む広範囲のモデルを評価するための自動採点パイプラインを開発する。blind-spots-benchでの分析により、クローズドソースのフロンティアモデルは、既存のベンチマークで同等の性能を示す場合でも、オープンウェイトモデルを約10%の差で大幅に上回ることが明らかになった。より詳細な分析では、単一のモデルが全てのタスクタイプで支配的であるわけではなく、一部のタスクは評価した全てのモデルにとって困難なままであることが示された。これらの結果は、現在の最新モデルにおける具体的な弱点を特定するための診断的ストレステストとしてのblind-spots-benchの価値を強調している。
English
Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs. These examples suggest that existing benchmarks may under-measure persistent blind spots in current systems. We introduce blind-spots-bench, a benchmark designed to expose such blind spots through tasks that appear simple for humans but remain challenging for modern AI. We collect raw questions from students in an AI course, clean and annotate them with structured reference solutions, and propose a task taxonomy tailored to the resulting dataset of 235 samples. We further develop an automated grading pipeline to evaluate a wide range of models, including open-weight and closed-source language, vision-language, and image-generation models. Our analysis on blind-spots-bench reveals that closed-source frontier models can substantially outperform open-weight models with even approx10% gap, even when they attain comparable performance on existing benchmarks. A more fine-grained analysis shows that no single model dominates across all task types, and that some tasks remain challenging for all evaluated models. These results highlight the value of blind-spots-bench as a diagnostic stress test for identifying concrete weaknesses in current modern models.