Blind-Spots-Bench: 멀티모달 모델의 맹점 평가
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
July 9, 2026
저자: Matteo Santelmo, Xiuying Wei, Israa Fakih, Felix Bauer, Juan Garcia Giraldo, Chengkun Li, Etienne Bamas, Emmanuel Abbé
cs.AI
초록
현대 AI 모델은 많은 기존 벤치마크에서 뛰어난 성능을 달성했지만, 여전히 인간이 거의 사소하다고 여기는 작업(예: 끈을 조작하거나 다섯 개의 다리를 가진 개를 그리기)에서는 실패한다. 이러한 사례는 기존 벤치마크가 현재 시스템의 지속적인 사각지대를 충분히 측정하지 못할 수 있음을 시사한다. 본 연구에서는 인간에게는 단순해 보이지만 현대 AI에는 여전히 어려운 작업을 통해 이러한 사각지대를 드러내도록 설계된 벤치마크인 blind-spots-bench를 소개한다. AI 수업의 학생들로부터 원시 질문을 수집하고, 이를 정리하여 구조화된 참조 해답과 함께 주석을 달았으며, 결과적으로 얻은 235개 샘플의 데이터셋에 맞춤화된 작업 분류 체계를 제안한다. 또한 오픈웨이트 모델과 폐쇄형 소스 언어, 비전-언어, 이미지 생성 모델을 포함한 다양한 모델을 평가하기 위한 자동 채점 파이프라인을 개발한다. blind-spots-bench에 대한 분석 결과, 폐쇄형 최첨단 모델이 오픈웨이트 모델보다 약 10%의 격차로 크게 우수한 성능을 보일 수 있으며, 이는 기존 벤치마크에서 유사한 성능을 보일 때조차 그러하다. 더 세분화된 분석은 어떤 단일 모델도 모든 작업 유형에서 지배적이지 않으며, 일부 작업은 평가된 모든 모델에 대해 여전히 어려움을 겪는다는 것을 보여준다. 이러한 결과는 blind-spots-bench가 현재 현대 모델의 구체적인 취약점을 식별하기 위한 진단적 스트레스 테스트로서의 가치를 강조한다.
English
Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs. These examples suggest that existing benchmarks may under-measure persistent blind spots in current systems. We introduce blind-spots-bench, a benchmark designed to expose such blind spots through tasks that appear simple for humans but remain challenging for modern AI. We collect raw questions from students in an AI course, clean and annotate them with structured reference solutions, and propose a task taxonomy tailored to the resulting dataset of 235 samples. We further develop an automated grading pipeline to evaluate a wide range of models, including open-weight and closed-source language, vision-language, and image-generation models. Our analysis on blind-spots-bench reveals that closed-source frontier models can substantially outperform open-weight models with even approx10% gap, even when they attain comparable performance on existing benchmarks. A more fine-grained analysis shows that no single model dominates across all task types, and that some tasks remain challenging for all evaluated models. These results highlight the value of blind-spots-bench as a diagnostic stress test for identifying concrete weaknesses in current modern models.