CLIP-CC-Bench: ビデオ言語モデルにおける段落レベルのビデオ記述の評価
CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models
August 5, 2026
著者: Mukhtiar Ali, Harsh Dubey, Sugam Mishra, Chulwoo Pack
cs.AI
要旨
映像言語モデルのベンチマーク評価は、これまで主に短いクリップと単文レベルの指標に焦点を当てており、現在のシステムが正確な長文・段落レベルの記述を生成できるかどうかは未解明のままである。本稿では、5時間分の映画コンテンツを90秒のクリップに分割し、各クリップに専門家が執筆した段落形式の参照記述を対応付けた、長文映像記述のための評価スイートCLIP-CC-Benchを導入する。本評価スイートは、5つの最先端LLMベースの埋め込みモデルのアンサンブルを採用して信頼性を高め、単一モデルのバイアスを軽減する。また、(i)粗粒度の意味マッチングと(ii)細粒度の意味マッチングという2つの相補的手法を適用し、モデルが生成した記述をCLIP-CC-Benchの参照記述と比較する。この枠組みを用いて、17の最先端映像言語モデルを評価し、ボルダ集約によるランキングとCLIP-CC-Benchにおける平均スコアの両方を報告する。さらに、判定者間一致度とブートストラップ法によるランキング安定性を通じて、本プロトコルの内部信頼性を定量化する。再現可能性を支援するため、標準化された評価スクリプト、モデル出力、集約ツールをhttps://github.com/Multimodal-Intelligence-Lab/CLIP-CC-Benchで公開する。CLIP-CC-Benchは、既存の短いクリップ向けおよびQA専用のベンチマークが残したギャップを埋め、長文映像記述のための実用的な評価枠組みを提供するものである。
English
Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate accurate long-form, paragraph-level descriptions. We introduce CLIP-CC-Bench, an evaluation suite for long-form video description built from 5 hours of movie content segmented into 90-second clips, each paired with an expert-written paragraph-style reference. The evaluation suite employs an ensemble of five state-of-the-art LLM-based embedding models to increase reliability and mitigate single-model bias, and applies two complementary methodologies: (i) coarse-grained semantic matching and (ii) fine-grained semantic matching to compare model-generated descriptions against CLIP-CC-Bench references. Using this framework, we evaluate 17 state-of-the-art video-language models and report both their Borda-aggregated rankings and their average scores on CLIP-CC-Bench. We further quantify the protocol's internal reliability through inter-judge agreement and bootstrap ranking stability. We release standardized evaluation scripts, model outputs, and aggregation tools at https://github.com/Multimodal-Intelligence-Lab/CLIP-CC-Bench to support reproducibility. CLIP-CC-Bench provides a practical evaluation framework for long-form video description, filling a gap left by existing short-clip and QA-only benchmarks.