KeyFrame-Compass: キーフレーム条件付きビデオ生成の包括的評価に向けて
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
July 15, 2026
著者: Yuqi Tang, Tengfei Liu, Yizheng Lai, Yuran Wang, Yang Shi, Wanshun Su, Zhuoran Zhang, Qixun Wang, Xiaohan Zhang, Xinlei Yu, Xuehai Bai, Xuanyu Zhu, Bohan Zeng, Bozhou Li, Shujie Li, Yifan Dai, Yujie Wei, Shixuan Liu, Haotian Wang, Jialu Chen, Yuanxing Zhang
cs.AI
要旨
ビデオ生成はますますキーフレームベースのワークフローに依存するようになっており、クリエイターは一連の参照画像を指定して生成を誘導する。近年のモデルは複数キーフレームによる条件付けをサポートしているものの、全体的なビデオ品質を維持しながら指定されたキーフレームを忠実に再現できるかは不明である。本稿では、キーフレーム条件付きビデオ生成を評価するための初の包括的ベンチマークであるKeyFrame-Compassを提案する。このベンチマークは、3つの応用ドメイン、2つのビデオ構造、2つのプロンプト粒度、2つの条件付け形式、4つのキーフレーム密度にわたる386の厳選サンプルを含み、多様な生成設定下での制御された分析を可能にする。さらに、キーフレームの実行と全体的なビデオ品質を共同で測定する自動評価フレームワークを導入する。具体的には、キーフレーム実行を、存在性、忠実性、時間的順序、局所性、持続性、一意性をカバーする6つの補完的メトリクスに分解し、一方で全体的なビデオ品質は、専門的な知覚モデルで拡張された証拠に基づくMLLM判断を通じて評価する。9つの代表的なビデオ生成システムに関する実験により、いくつかの根本的な限界が明らかになった。現在のモデルは、忠実なキーフレーム実行と自然なビデオ合成の間に明確なトレードオフを示している。さらに、キーフレームの制約が密になるにつれて性能は低下し、ほとんどのオープンソースモデルはストーリーボードグリッド入力を時間的に順序付けられたキーフレームシーケンスとして解釈できない。
English
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.