ChatPaper.aiChatPaper

PerceptionBench:评估多模态大语言模型中的原子视觉感知

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

July 27, 2026
作者: Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang, Jia Li, Haoning Wu, Yuhao Dong, Zuhao Yang, Jinguo Zhu, Haoyu Lu, Zijia Zhao, Tongtian Yue, Zhangyang Qi, Junwei Yang, Mengfan Dong, Peizhou Cao, Chenzhuang Du, Zaida Zhou, Haotian Yao, Hao Yang, Hongcheng Gao, Lin Sui, Weihong Li, Xinxing Zu, Jia Chen, Yao Wang, Xiaoxue Wu, Yalin Wang, Y. Charles, Yiping Bao, Yangyang Liu, Zhiqi Huang, Xinyu Zhou
cs.AI

摘要

我们提出了PerceptionBench——一个专门用于评估多模态大语言模型(MLLMs)原子视觉感知能力的基准。现有基准往往无法隔离感知能力:整体评估将感知错误与推理或领域知识的失败混为一谈,而应用驱动的基准仅覆盖由启发式设计所塑造的狭窄、碎片化的领域。为克服这些局限,PerceptionBench采用自下而上的方法:通过诊断前沿MLLMs在42个现有基准中最早出现的失败点,我们构建了一个错误分类体系,其中感知分支定义了十种原子感知能力。在该分类体系的指导下,我们构建了3000个经过验证的问题,每个问题答案简短明确,且仅隔离单一能力,其难度源于感知而非推理或知识。在16个前沿MLLMs上的基准测试结果表明,原子感知能力基本尚未解决——没有模型达到60%的准确率,与感知相关的幻觉是平均表现最弱的能力,而相似的总分背后隐藏着截然不同的能力分布。因此,PerceptionBench为衡量和诊断MLLMs的视觉感知边界提供了能力层面的标准。
English
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.