ChatPaper.aiChatPaper

關鍵幀指南針:邁向關鍵幀條件下影片生成的全面評測

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

July 15, 2026
作者: Yuqi Tang, Tengfei Liu, Yizheng Lai, Yuran Wang, Yang Shi, Wanshun Su, Zhuoran Zhang, Qixun Wang, Xiaohan Zhang, Xinlei Yu, Xuehai Bai, Xuanyu Zhu, Bohan Zeng, Bozhou Li, Shujie Li, Yifan Dai, Yujie Wei, Shixuan Liu, Haotian Wang, Jialu Chen, Yuanxing Zhang
cs.AI

摘要

影片生成日益依賴以關鍵幀為基礎的流程,創作者指定一系列參考影像來引導生成過程。儘管近期模型已支援多重關鍵幀條件設定,但它們能否在維持整體影片品質的同時忠實重現指定的關鍵幀,仍有待釐清。我們提出 KeyFrame-Compass,這是首個專門評估關鍵幀條件影片生成的綜合基準。該基準包含 386 個精心挑選的樣本,涵蓋三個應用領域、兩種影片結構、兩種提示粒度、兩種條件設定格式及四種關鍵幀密度,能在多樣生成設定下進行受控分析。我們進一步引入一套自動化評估框架,同時衡量關鍵幀執行度與整體影片品質。具體而言,我們將關鍵幀執行度拆解為六項互補指標:存在性、保真度、時序正確性、定位準確性、持續性及唯一性;同時透過具備證據基礎的多模態大語言模型判斷,並輔以專門的感知模型,來評估整體影片品質。針對九個代表性影片生成系統的實驗揭示了數項根本限制。當前模型在忠實執行關鍵幀與自然影片合成之間存在明顯的權衡取捨。隨著關鍵幀約束密度增加,模型效能進一步下滑,且多數開源模型亦無法將分鏡格輸入解讀為時序排列的關鍵幀序列。
English
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.