ChatPaper.aiChatPaper

PerceptionBench: 다중 모달 대규모 언어 모델에서의 원자적 시각 인지 평가

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

July 27, 2026
저자: Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang, Jia Li, Haoning Wu, Yuhao Dong, Zuhao Yang, Jinguo Zhu, Haoyu Lu, Zijia Zhao, Tongtian Yue, Zhangyang Qi, Junwei Yang, Mengfan Dong, Peizhou Cao, Chenzhuang Du, Zaida Zhou, Haotian Yao, Hao Yang, Hongcheng Gao, Lin Sui, Weihong Li, Xinxing Zu, Jia Chen, Yao Wang, Xiaoxue Wu, Yalin Wang, Y. Charles, Yiping Bao, Yangyang Liu, Zhiqi Huang, Xinyu Zhou
cs.AI

초록

본 논문에서는 다중 모달 대규모 언어 모델(MLLM)의 원자적 시각 인지 능력을 평가하기 위해 특별히 설계된 벤치마크인 PerceptionBench를 소개한다. 기존 벤치마크는 종종 인지를 고립적으로 평가하지 못한다. 즉, 전체론적 평가는 인지 오류를 추론이나 도메인 지식의 실패와 혼동하며, 애플리케이션 중심 벤치마크는 휴리스틱 설계에 의해 형성된 좁고 단편적인 영역만을 다룬다. 이러한 한계를 해결하기 위해 PerceptionBench는 상향식 접근 방식을 채택한다. 즉, 42개의 기존 벤치마크에 걸쳐 선도적 MLLM 응답에서 가장 초기에 발생하는 실패 지점을 진단함으로써 오류 분류 체계를 구축하고, 이 체계의 인지 분기는 10가지 원자적 인지 능력을 정의한다. 이 분류 체계를 바탕으로, 각 능력을 고립적으로 평가하는 짧고 명확한 답변을 가진 3,000개의 검증된 질문을 구축하였으며, 난이도는 추론이나 지식이 아닌 인지에서 비롯된다. 16개의 선도적 MLLM에 대한 벤치마크 결과는 원자적 인지가 여전히 해결되지 않은 문제임을 보여준다. 어떤 모델도 60% 정확도에 도달하지 못했으며, 인지 관련 환각이 평균적으로 가장 약한 능력이고, 유사한 전체 점수가 급격히 다른 능력 프로파일을 숨기고 있다. 이에 따라 PerceptionBench는 MLLM의 시각 인지 경계를 측정하고 진단하기 위한 능력 수준의 표준을 제공한다.
English
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.