ChatPaper.aiChatPaper

CAPEval: 이해와 생성에 걸친 분리된 캡션 평가

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

August 3, 2026
저자: Zhipeng Liu, Haochen Wang, Zhaoxiang Zhang
cs.AI

초록

캡션은 다중 모달 이해와 텍스트-이미지 생성 모두에서 일차적인 지도 신호로 기능한다. 그러나 기존 평가는 캡션 품질을 단일 스칼라 목표로 취급하여, (1) 캡션이 얼마나 많은 시각 정보를 포함하는지와 (2) 이미지가 자신이 명시하는 주장을 얼마나 신뢰성 있게 뒷받침하는지라는 두 가지 별개의 속성을 혼동한다. 이에 우리는 인간이 작성한 정답 캡션과 인간이 검증한 원자적 체크리스트 항목을 갖춘 분리형 캡션 평가 벤치마크인 CAPEval(Coverage And Precision Evaluation)을 설계한다. 구체적으로, CAPEval은 캡션 품질을 Coverage(커버리지)와 Precision(정밀도)으로 분해한다. 전자는 캡션이 정답 사실적 내용을 얼마나 철저히 포함하는지를 정량화하며, 후자는 캡션에 표현된 모든 주장의 사실 정확도 비율을 반영한다. 우리는 10개의 캡셔너를 선정하고, 캡션 소스가 유일한 변수인 네 가지 모델 계열에 걸친 통제된 다운스트림 엔드투엔드 실험을 추가로 수행한다. 실험 결과, 우리는 일관된 작업 의존적 분리를 발견한다. 즉 Coverage는 이해 성능과 더 강한 상관관계를 보이는 반면, Precision은 생성 성능의 지배적인 예측 변수로 작용한다. 이러한 분리형 평가 패러다임은 캡션 품질에 대한 더 세밀한 진단을 제공할 뿐만 아니라, 다양한 다운스트림 작업에 맞춰 캡셔너를 선택하고 최적화하기 위한 실행 가능한 지침을 제공한다.
English
Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information a caption covers and (2) how reliably the image supports its stated claims. To this end, we design a decoupled caption evaluation benchmark, CAPEval (Coverage And Precision Evaluation), with human-written ground-truth captions and human-verified atomic checklist items. Specifically, CAPEval decomposes caption quality into Coverage and Precision. The former quantifies how thoroughly a caption covers ground-truth factual content, while the latter reflects the factual correctness rate of all claims expressed in the caption. We select 10 captioners and further conduct controlled downstream end-to-end experiments with them from four model families, where the caption source is the only variable. Empirically, we find a consistent task-dependent dissociation: Coverage serves as the stronger correlate for understanding performance, whereas Precision acts as the dominant predictor for generation performance. This decoupled evaluation paradigm not only delivers a more fine-grained diagnosis of caption quality, but also offers actionable guidance for selecting and optimizing captioners tailored to different downstream tasks.