ChatPaper.aiChatPaper

FilmBench:電影級影片生成的電影級基準

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

July 27, 2026
作者: Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi, Fei Ding, Weixu Qiao, Jinlin Wang, Xiaotong Lv, Peng Han, Zimeng Li, Fanshu Ding, Yushu Wang, Han Wu, Jingjing Chen, Chongxiao Wang, Yanhao Wu, Chenglong Huang, Xiaoqian Zhu, Jie Tian, Hua Li, Jingjing Fan, Mingshuang Tang, Zhong Li, Hengxia Qiang, Weibin Chen, Jinyang Zhen, Bing Zhao, Lin Qu, Jing Li, Hu Wei
cs.AI

摘要

視訊生成技術的進展持續縮小人工智慧生成影片與專業製作影片之間的視覺差距,然而多數基準測試仍從網路來源或大型語言模型模板提取提示,並以未經訓練的通用多模態模型進行評分。更根本的問題在於,這些評測的分類體系仍相當粗淺(整體視覺品質、粗略文字對齊與時間連續性),而非專業電影製作與評判所依據的「電影語言」準則,因此它們僅評估影片的基本合理性,而非電影級別的工藝。我們提出FilmBench——一個根植於電影學院傳統專業電影語言、並與北京電影學院及互晶數位媒體娛樂集團電影製片廠的導演與教師共同開發的文字生成影片(T2V)與參考影片生成(R2V)基準。該基準基於三項選擇。第一,提示是從涵蓋20種電影類型的獲獎影片片段中逆向工程生成,並由專業導演挑選,因此每個提示都錨定於經過驗證的真人拍攝參考;提示遵循實際鏡頭清單,且大多數腳本包含多個鏡頭(1,169個提示中有1,056個為多鏡頭),有別於以往的單一鏡頭基準。第二,評測遵循一個三層級的電影分類體系,包含3個軸向、12個組成部分及35個(T2V)+3個(僅R2V)子指標。第三,我們開發了一個內部的專家級自動評測代理,並開源其核心電影語言運算子套件(FilmOps)。在對領先的視訊生成模型(9個T2V模型、7個R2V模型)進行基準測試時,該評測器在模型層級上重現了人類模型排序,其斯皮爾曼相關係數分別為0.95(T2V)與0.96(R2V)。分數遠低於先前的網路風格基準,存在兩個持續的差距:動態美學方面,以及從單鏡頭到多鏡頭的性能顯著下降,且此下降在較弱模型中更為明顯。
English
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman ho = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.