ChatPaper.aiChatPaper

PerceptionBench: マルチモーダル大規模言語モデルにおけるアトミックな視覚知覚の評価

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

July 27, 2026
著者: Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang, Jia Li, Haoning Wu, Yuhao Dong, Zuhao Yang, Jinguo Zhu, Haoyu Lu, Zijia Zhao, Tongtian Yue, Zhangyang Qi, Junwei Yang, Mengfan Dong, Peizhou Cao, Chenzhuang Du, Zaida Zhou, Haotian Yao, Hao Yang, Hongcheng Gao, Lin Sui, Weihong Li, Xinxing Zu, Jia Chen, Yao Wang, Xiaoxue Wu, Yalin Wang, Y. Charles, Yiping Bao, Yangyang Liu, Zhiqi Huang, Xinyu Zhou
cs.AI

要旨

我々は、マルチモーダル大規模言語モデル(MLLM)の原子視覚知覚能力を評価するために特別に設計されたベンチマーク「PerceptionBench」を紹介する。既存のベンチマークはしばしば知覚を分離することに失敗している。すなわち、全体的な評価は知覚エラーと推論やドメイン知識における失敗を混同し、アプリケーション主導のベンチマークはヒューリスティックな設計によって形成された狭く断片的な領域しかカバーしていない。これらの限界に対処するため、PerceptionBenchはボトムアップアプローチを採用する。42の既存ベンチマークにわたる最先端MLLMの応答における最も初期の失敗点を診断することにより、エラータクソノミを構築し、その知覚枝が10の原子知覚能力を定義する。このタクソノミに導かれて、我々は短く曖昧でない回答を持つ3,000の検証済み質問を構築し、それぞれが単一の能力を分離し、難易度は推論や知識ではなく知覚に起因する。16の最先端MLLMにわたるベンチマーク結果は、原子知覚が依然としてほとんど解決されていないことを明らかにしている。どのモデルも60%の精度に達しておらず、知覚関連の幻覚は平均して最も弱い能力であり、類似した全体スコアの背後に著しく異なる能力プロファイルが隠れている。したがって、PerceptionBenchはMLLMの視覚知覚の限界を測定・診断するための能力レベルの基準を提供する。
English
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.