ChatPaper.aiChatPaper

盲点基准:评估多模态模型中的盲点

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

July 9, 2026
作者: Matteo Santelmo, Xiuying Wei, Israa Fakih, Felix Bauer, Juan Garcia Giraldo, Chengkun Li, Etienne Bamas, Emmanuel Abbé
cs.AI

摘要

现代AI模型在许多既定基准测试中表现优异,但在人类几乎能轻易完成的任务上仍然失败,例如操控一根绳子或画一只五条腿的狗。这些例子表明,现有基准测试可能低估了当前系统中持续存在的盲点。我们推出blind-spots-bench——一个旨在通过人类看似简单但对现代AI仍具挑战性的任务来暴露此类盲点的基准测试。我们从AI课程的学生中收集原始问题,进行清洗并添加结构化参考解决方案的注释,针对由此产生的235个样本数据集提出了一套任务分类体系。我们进一步开发了自动化评分流程,用于评估包括开源权重模型、闭源模型、视觉语言模型以及图像生成模型在内的广泛模型。我们对blind-spots-bench的分析发现,闭源前沿模型能够显著超越开源权重模型,差距甚至达到约10%,即便它们在现有基准测试中表现相当。更细粒度的分析表明,没有单一模型在所有任务类型上占据主导地位,且某些任务对所有评估模型而言依然充满挑战。这些结果凸显了blind-spots-bench作为诊断性压力测试的价值,有助于识别当前现代模型的具体弱点。
English
Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs. These examples suggest that existing benchmarks may under-measure persistent blind spots in current systems. We introduce blind-spots-bench, a benchmark designed to expose such blind spots through tasks that appear simple for humans but remain challenging for modern AI. We collect raw questions from students in an AI course, clean and annotate them with structured reference solutions, and propose a task taxonomy tailored to the resulting dataset of 235 samples. We further develop an automated grading pipeline to evaluate a wide range of models, including open-weight and closed-source language, vision-language, and image-generation models. Our analysis on blind-spots-bench reveals that closed-source frontier models can substantially outperform open-weight models with even approx10% gap, even when they attain comparable performance on existing benchmarks. A more fine-grained analysis shows that no single model dominates across all task types, and that some tasks remain challenging for all evaluated models. These results highlight the value of blind-spots-bench as a diagnostic stress test for identifying concrete weaknesses in current modern models.