CLIP-CC-Bench:評估影片語言模型中的段落級影片描述
CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models
August 5, 2026
作者: Mukhtiar Ali, Harsh Dubey, Sugam Mishra, Chulwoo Pack
cs.AI
摘要
影片-語言模型的基準測試大多聚焦於短片與單句指標,尚待探討現有系統能否生成準確的長篇段落級描述。我們提出 CLIP-CC-Bench,這是一個用於長篇影片描述的評估套件,從 5 小時的電影內容建構而成,分割為 90 秒的片段,每個片段搭配專家撰寫的段落式參考描述。該評估套件採用五個最先進的基於大型語言模型(LLM)的嵌入模型所組成的集成,以提高可靠性並減輕單一模型的偏誤,並應用兩種互補方法:(i) 粗粒度語意比對與 (ii) 細粒度語意比對,以比較模型生成的描述與 CLIP-CC-Bench 參考描述。利用此框架,我們評估了 17 個最先進的影片-語言模型,並報告它們在 CLIP-CC-Bench 上的 Borda 聚合排名與平均分數。我們進一步透過評分者間一致性與拔靴法排名穩定性來量化該協定的內部信度。我們在 https://github.com/Multimodal-Intelligence-Lab/CLIP-CC-Bench 發布標準化評估腳本、模型輸出與聚合工具,以支持可重現性。CLIP-CC-Bench 為長篇影片描述提供了實用的評估框架,填補了現有短片與僅限問答基準所留下的缺口。
English
Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate accurate long-form, paragraph-level descriptions. We introduce CLIP-CC-Bench, an evaluation suite for long-form video description built from 5 hours of movie content segmented into 90-second clips, each paired with an expert-written paragraph-style reference. The evaluation suite employs an ensemble of five state-of-the-art LLM-based embedding models to increase reliability and mitigate single-model bias, and applies two complementary methodologies: (i) coarse-grained semantic matching and (ii) fine-grained semantic matching to compare model-generated descriptions against CLIP-CC-Bench references. Using this framework, we evaluate 17 state-of-the-art video-language models and report both their Borda-aggregated rankings and their average scores on CLIP-CC-Bench. We further quantify the protocol's internal reliability through inter-judge agreement and bootstrap ranking stability. We release standardized evaluation scripts, model outputs, and aggregation tools at https://github.com/Multimodal-Intelligence-Lab/CLIP-CC-Bench to support reproducibility. CLIP-CC-Bench provides a practical evaluation framework for long-form video description, filling a gap left by existing short-clip and QA-only benchmarks.