ChatPaper.aiChatPaper

CAPEval: 理解と生成にわたる分離型キャプション評価

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

August 3, 2026
著者: Zhipeng Liu, Haochen Wang, Zhaoxiang Zhang
cs.AI

要旨

キャプションは、マルチモーダル理解とテキストから画像への生成の両方において、主要な教師信号として機能する。しかしながら、従来の評価ではキャプション品質を単一のスカラー目的として扱っており、それによって以下の2つの異なる特性を混同している。(1) キャプションがどの程度の視覚情報を網羅しているか、そして (2) 画像がキャプション内の明示的な主張をどの程度確実に裏付けているか。この目的のため、我々は、人間が作成したグラウンドトゥルースキャプションと人間が検証した原子的なチェックリスト項目を用いた、分離型キャプション評価ベンチマークであるCAPEval(Coverage And Precision Evaluation)を設計する。具体的には、CAPEvalはキャプション品質をカバレッジ(網羅性)とプレシジョン(正確性)に分解する。前者はキャプションがグラウンドトゥルースの事実内容をどの程度徹底的に網羅しているかを定量化し、後者はキャプション内で表現されたすべての主張の事実正確性率を反映する。我々は10種類のキャプション生成器を選択し、さらに4つのモデルファミリーを用いて、キャプションソースのみを変数とした統制された下流エンドツーエンド実験を実施する。経験的に、我々は一貫したタスク依存の解離を発見する。すなわち、カバレッジは理解性能に対してより強い相関因子として機能し、一方プレシジョンは生成性能に対する支配的な予測因子として機能する。この分離型評価パラダイムは、キャプション品質のよりきめ細かい診断を提供するだけでなく、異なる下流タスクに適合したキャプション生成器の選択と最適化のための実行可能な指針も提供する。
English
Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information a caption covers and (2) how reliably the image supports its stated claims. To this end, we design a decoupled caption evaluation benchmark, CAPEval (Coverage And Precision Evaluation), with human-written ground-truth captions and human-verified atomic checklist items. Specifically, CAPEval decomposes caption quality into Coverage and Precision. The former quantifies how thoroughly a caption covers ground-truth factual content, while the latter reflects the factual correctness rate of all claims expressed in the caption. We select 10 captioners and further conduct controlled downstream end-to-end experiments with them from four model families, where the caption source is the only variable. Empirically, we find a consistent task-dependent dissociation: Coverage serves as the stronger correlate for understanding performance, whereas Precision acts as the dominant predictor for generation performance. This decoupled evaluation paradigm not only delivers a more fine-grained diagnosis of caption quality, but also offers actionable guidance for selecting and optimizing captioners tailored to different downstream tasks.