ChatPaper.aiChatPaper

KeyFrame-Compass: 키프레임 조건부 비디오 생성의 포괄적 평가를 향하여

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

July 15, 2026
저자: Yuqi Tang, Tengfei Liu, Yizheng Lai, Yuran Wang, Yang Shi, Wanshun Su, Zhuoran Zhang, Qixun Wang, Xiaohan Zhang, Xinlei Yu, Xuehai Bai, Xuanyu Zhu, Bohan Zeng, Bozhou Li, Shujie Li, Yifan Dai, Yujie Wei, Shixuan Liu, Haotian Wang, Jialu Chen, Yuanxing Zhang
cs.AI

초록

비디오 생성은 점점 더 키프레임 기반 워크플로에 의존하고 있으며, 여기서 창작자는 일련의 참조 이미지를 지정하여 생성을 안내한다. 최근 모델들이 다중 키프레임 조건 설정을 지원하지만, 전체 비디오 품질을 유지하면서 지정된 키프레임을 충실히 재현할 수 있는지는 여전히 불분명하다. 본 논문에서는 키프레임 조건 비디오 생성을 평가하기 위한 최초의 포괄적 벤치마크인 KeyFrame-Compass를 제시한다. 이 벤치마크는 세 가지 응용 도메인, 두 가지 비디오 구조, 두 가지 프롬프트 세분성, 두 가지 조건 설정 형식, 네 가지 키프레임 밀도에 걸쳐 엄선된 386개의 샘플을 포함하여 다양한 생성 설정에서 제어된 분석을 가능하게 한다. 또한 키프레임 실행과 전체 비디오 품질을 동시에 측정하는 자동 평가 프레임워크를 도입한다. 구체적으로 키프레임 실행을 존재 여부, 충실도, 시간적 순서, 위치화, 지속성, 고유성을 포괄하는 여섯 가지 상호 보완적 지표로 분해하고, 특화된 인식 모델로 강화된 증거 기반 MLLM 판단을 통해 전체 비디오 품질을 평가한다. 9개의 대표적인 비디오 생성 시스템에 대한 실험은 몇 가지 근본적인 한계를 드러낸다. 현재 모델들은 충실한 키프레임 실행과 자연스러운 비디오 합성 사이에서 명확한 트레이드오프를 보인다. 키프레임 제약이 밀집할수록 성능이 더욱 저하되며, 대부분의 오픈소스 모델은 스토리보드 그리드 입력을 시간적으로 정렬된 키프레임 시퀀스로 해석하는 데 실패한다.
English
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.