PerceptionBench:評估多模態大型語言模型中的原子視覺感知
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
July 27, 2026
作者: Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang, Jia Li, Haoning Wu, Yuhao Dong, Zuhao Yang, Jinguo Zhu, Haoyu Lu, Zijia Zhao, Tongtian Yue, Zhangyang Qi, Junwei Yang, Mengfan Dong, Peizhou Cao, Chenzhuang Du, Zaida Zhou, Haotian Yao, Hao Yang, Hongcheng Gao, Lin Sui, Weihong Li, Xinxing Zu, Jia Chen, Yao Wang, Xiaoxue Wu, Yalin Wang, Y. Charles, Yiping Bao, Yangyang Liu, Zhiqi Huang, Xinyu Zhou
cs.AI
摘要
我們提出了PerceptionBench,這是一個專門設計用來評估多模態大型語言模型(MLLMs)原子視覺感知能力的基準測試。現有的基準測試往往無法孤立地評估感知:整體評估將感知錯誤與推理或領域知識的失敗混為一談,而應用導向的基準測試僅涵蓋由啟發式設計形成的狹窄且零碎的領域。為了解決這些限制,PerceptionBench採用自下而上的方法:透過診斷前沿MLLMs在42個現有基準測試中回應的最早失敗點,我們建立了一個錯誤分類體系,其感知分支定義了十項原子感知能力。在此分類體系的引導下,我們建構了3,000個經過驗證的問題,每個問題具有簡短明確的答案,且各自獨立測量單一能力,其難度源自感知而非推理或知識。在十六個前沿MLLMs上的基準測試結果顯示,原子感知在很大程度上仍未獲得解決——沒有模型達到60%的準確率,與感知相關的幻覺是平均表現最弱的能力,而相似的整体分數掩蓋了截然不同的能力輪廓。因此,PerceptionBench為衡量和診斷MLLMs的視覺感知邊界提供了能力層面的標準。
English
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.