ChatPaper.aiChatPaper

CLIP-CC-Bench:评估视频-语言模型中的段落级视频描述

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models

August 5, 2026
作者: Mukhtiar Ali, Harsh Dubey, Sugam Mishra, Chulwoo Pack
cs.AI

摘要

视频-语言模型的基准评测主要集中于短视频片段和单句指标,当前系统能否生成准确的长篇、段落级描述这一问题仍未得到解答。我们提出了 CLIP-CC-Bench,这是一个用于长篇视频描述的评测套件,基于5小时的电影内容构建,并将其分割为90秒的片段,每个片段配有专家撰写的段落式参考描述。该评测套件采用五种最先进的基于大语言模型的嵌入模型组成的集成,以提高可靠性并减轻单一模型的偏差,同时应用两种互补的方法论:(i) 粗粒度语义匹配和 (ii) 细粒度语义匹配,将模型生成的描述与 CLIP-CC-Bench 的参考描述进行比较。利用这一框架,我们评估了17个最先进的视频-语言模型,并报告了它们在 CLIP-CC-Bench 上的波达聚合排名及平均得分。我们还通过评估者间一致性和自助法排名稳定性,进一步量化了该评测协议的内部可靠性。我们在 https://github.com/Multimodal-Intelligence-Lab/CLIP-CC-Bench 上发布标准化的评测脚本、模型输出和聚合工具,以支持可复现性。CLIP-CC-Bench 为长篇视频描述提供了一个实用的评测框架,填补了现有短片段和纯问答基准所留下的空白。
English
Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate accurate long-form, paragraph-level descriptions. We introduce CLIP-CC-Bench, an evaluation suite for long-form video description built from 5 hours of movie content segmented into 90-second clips, each paired with an expert-written paragraph-style reference. The evaluation suite employs an ensemble of five state-of-the-art LLM-based embedding models to increase reliability and mitigate single-model bias, and applies two complementary methodologies: (i) coarse-grained semantic matching and (ii) fine-grained semantic matching to compare model-generated descriptions against CLIP-CC-Bench references. Using this framework, we evaluate 17 state-of-the-art video-language models and report both their Borda-aggregated rankings and their average scores on CLIP-CC-Bench. We further quantify the protocol's internal reliability through inter-judge agreement and bootstrap ranking stability. We release standardized evaluation scripts, model outputs, and aggregation tools at https://github.com/Multimodal-Intelligence-Lab/CLIP-CC-Bench to support reproducibility. CLIP-CC-Bench provides a practical evaluation framework for long-form video description, filling a gap left by existing short-clip and QA-only benchmarks.