ChatPaper.aiChatPaper

CAPEval:一种跨理解与生成任务的解耦式图像描述评估方法

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

August 3, 2026
作者: Zhipeng Liu, Haochen Wang, Zhaoxiang Zhang
cs.AI

摘要

标题(caption)在图像理解与文生图生成任务中均被用作主要的监督信号。然而,现有评测往往将描述质量视为单一标量目标,从而混淆了两个截然不同的属性:(1)描述覆盖了多少视觉信息;(2)图像能在多大程度上可靠地支持描述中所陈述的内容。为此,我们设计了一个解耦式描述评测基准——CAPEval(覆盖度与精确度评测,Coverage And Precision Evaluation),该基准包含人工撰写的真值描述以及经人工核验的原子化检查表条目。具体而言,CAPEval 将描述质量分解为覆盖度(Coverage)与精确度(Precision)两个维度:前者衡量描述对真值事实内容的覆盖程度,后者反映描述中所有陈述的事实正确率。我们选取了 10 个描述生成器,并进一步在来自四个模型家族的描述生成器上进行了受控的下游端到端实验,其中描述来源是唯一的变量。实验结果表明,我们观察到一种一致的、随任务而异的分离现象:覆盖度与理解性能的相关性更强,而精确度则是生成性能的主导预测因子。这种解耦式评测范式不仅能够对描述质量提供更细粒度的诊断,还能为针对不同下游任务选择与优化描述生成器提供可操作的指导。
English
Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information a caption covers and (2) how reliably the image supports its stated claims. To this end, we design a decoupled caption evaluation benchmark, CAPEval (Coverage And Precision Evaluation), with human-written ground-truth captions and human-verified atomic checklist items. Specifically, CAPEval decomposes caption quality into Coverage and Precision. The former quantifies how thoroughly a caption covers ground-truth factual content, while the latter reflects the factual correctness rate of all claims expressed in the caption. We select 10 captioners and further conduct controlled downstream end-to-end experiments with them from four model families, where the caption source is the only variable. Empirically, we find a consistent task-dependent dissociation: Coverage serves as the stronger correlate for understanding performance, whereas Precision acts as the dominant predictor for generation performance. This decoupled evaluation paradigm not only delivers a more fine-grained diagnosis of caption quality, but also offers actionable guidance for selecting and optimizing captioners tailored to different downstream tasks.