FilmBench:面向电影感视频生成的电影级基准
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
July 27, 2026
作者: Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi, Fei Ding, Weixu Qiao, Jinlin Wang, Xiaotong Lv, Peng Han, Zimeng Li, Fanshu Ding, Yushu Wang, Han Wu, Jingjing Chen, Chongxiao Wang, Yanhao Wu, Chenglong Huang, Xiaoqian Zhu, Jie Tian, Hua Li, Jingjing Fan, Mingshuang Tang, Zhong Li, Hengxia Qiang, Weibin Chen, Jinyang Zhen, Bing Zhao, Lin Qu, Jing Li, Hu Wei
cs.AI
摘要
视频生成技术的进步不断缩小着人工智能生成画面与专业制作镜头之间的视觉差距,然而,多数基准测试仍从网络来源或大语言模型模板中抽取提示词,并使用未经训练的通用多模态模型进行评分。更根本的是,其评估分类体系仍停留在初级水平(整体视觉质量、粗略的文本对齐和时间平滑度),而非电影实际制作与评判所依据的专业电影语言标准,因此它们评估的是基础视频合理性而非电影级别的工艺。我们提出FilmBench,一个基于电影学院传统专业电影语言的文本到视频(T2V)及参考视频到视频(R2V)基准,由北京电影学院与互境数字传媒娱乐集团电影工作室的导演和教授共同开发。它基于三个设计选择。第一,提示词由获奖影片片段逆向推导而来,涵盖20种电影类型并由专业导演挑选,因此每个提示词都基于经过验证的真人实拍参考;提示词遵循真实的镜头列表,且多数描述多个镜头(1,169条提示词中有1,056条为多镜头),这与先前仅单镜头的基准不同。第二,评估采用三级电影分类体系,包含3个轴、12个组件以及35个(T2V)加3个(仅R2V)子指标。第三,我们开发了内部专家级自动评估代理,并开源其核心电影语言算子套件(FilmOps)。通过对主流视频生成模型(9个T2V模型,7个R2V模型)进行基准测试,该评估器在模型级别复现了人类排名,Spearman相关系数ρ=0.95(T2V)和0.96(R2V)。得分远低于先前的网络风格基准,在动态美学方面存在两个一致差距,且从单镜头到多镜头的性能显著下降,较弱模型的下滑幅度更大。
English
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman ho = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.