ChatPaper.aiChatPaper

盲點基準:評估多模態模型中的盲點

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

July 9, 2026
作者: Matteo Santelmo, Xiuying Wei, Israa Fakih, Felix Bauer, Juan Garcia Giraldo, Chengkun Li, Etienne Bamas, Emmanuel Abbé
cs.AI

摘要

現代AI模型在許多既定基準測試中表現出色,但在人類幾乎視為簡單的任務上仍會失敗,例如操控一條繩子或畫一隻五條腿的狗。這些案例顯示,現有基準可能低估了當前系統中持續存在的盲點。我們提出 blind-spots-bench,這是一個旨在透過對人類看似簡單但對現代AI仍具挑戰性的任務來揭露此類盲點的基準測試。我們從AI課程的學生中收集原始問題,進行清理並加上結構化參考解答的註解,再根據由此產生的235個樣本資料集提出一套任務分類法。我們進一步開發自動評分流程,以評估多種模型,包括開放權重與閉源的語言、視覺語言及影像生成模型。我們在 blind-spots-bench 上的分析顯示,即使閉源前沿模型在現有基準測試上與開放權重模型表現相當,前者仍可顯著超越後者,差距甚至約達10%。更細粒度的分析表明,沒有任何單一模型在所有任務類型中佔據主導地位,且部分任務對所有受評模型仍具挑戰性。這些結果凸顯了 blind-spots-bench 作為診斷壓力測試的價值,有助於識別當前現代模型中的具體弱點。
English
Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs. These examples suggest that existing benchmarks may under-measure persistent blind spots in current systems. We introduce blind-spots-bench, a benchmark designed to expose such blind spots through tasks that appear simple for humans but remain challenging for modern AI. We collect raw questions from students in an AI course, clean and annotate them with structured reference solutions, and propose a task taxonomy tailored to the resulting dataset of 235 samples. We further develop an automated grading pipeline to evaluate a wide range of models, including open-weight and closed-source language, vision-language, and image-generation models. Our analysis on blind-spots-bench reveals that closed-source frontier models can substantially outperform open-weight models with even approx10% gap, even when they attain comparable performance on existing benchmarks. A more fine-grained analysis shows that no single model dominates across all task types, and that some tasks remain challenging for all evaluated models. These results highlight the value of blind-spots-bench as a diagnostic stress test for identifying concrete weaknesses in current modern models.