KeyFrame-Compass:面向关键帧条件视频生成的全面评估
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
July 15, 2026
作者: Yuqi Tang, Tengfei Liu, Yizheng Lai, Yuran Wang, Yang Shi, Wanshun Su, Zhuoran Zhang, Qixun Wang, Xiaohan Zhang, Xinlei Yu, Xuehai Bai, Xuanyu Zhu, Bohan Zeng, Bozhou Li, Shujie Li, Yifan Dai, Yujie Wei, Shixuan Liu, Haotian Wang, Jialu Chen, Yuanxing Zhang
cs.AI
摘要
视频生成日益依赖基于关键帧的工作流程,创作者通过指定一系列参考图像来引导生成过程。尽管现有模型已支持多关键帧条件控制,但其能否在保持整体视频质量的同时忠实地复现指定关键帧仍不明确。我们提出KeyFrame-Compass,这是首个用于评估关键帧条件视频生成的综合性基准。该基准包含386个精心筛选的样本,覆盖三个应用领域、两种视频结构、两种提示粒度、两种条件格式及四种关键帧密度,可在多样化生成场景下实现受控分析。我们进一步引入自动化评估框架,同步测量关键帧执行度与整体视频质量。具体而言,我们将关键帧执行度分解为六个互补指标,涵盖存在性、保真度、时间顺序、定位精度、持续性与唯一性;同时通过基于证据的多模态大语言模型(MLLM)判断,并辅以专用感知模型,评估整体视频质量。对九种代表性视频生成系统的实验揭示了若干根本性局限:现有模型在忠实执行关键帧与自然视频合成之间呈现明显权衡;随着关键帧约束密度增加,模型性能进一步下降;多数开源模型无法将故事板网格输入正确解读为时序关键帧序列。
English
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.