ChatPaper.aiChatPaper

FilmBench: 映画級のシネマティック動画生成のためのベンチマーク

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

July 27, 2026
著者: Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi, Fei Ding, Weixu Qiao, Jinlin Wang, Xiaotong Lv, Peng Han, Zimeng Li, Fanshu Ding, Yushu Wang, Han Wu, Jingjing Chen, Chongxiao Wang, Yanhao Wu, Chenglong Huang, Xiaoqian Zhu, Jie Tian, Hua Li, Jingjing Fan, Mingshuang Tang, Zhong Li, Hengxia Qiang, Weibin Chen, Jinyang Zhen, Bing Zhao, Lin Qu, Jing Li, Hu Wei
cs.AI

要旨

動画生成の進歩により、AI生成映像とプロが制作した映像との視覚的ギャップは縮まり続けているが、現在のベンチマークのほとんどは依然としてWebソースやLLMテンプレートからプロンプトを抽出し、訓練されていない汎用のマルチモーダルモデルでスコアリングしている。より根本的には、その評価分類法は初歩的(全体的な視覚品質、大まかなテキストとの整合性、時間的滑らかさ)であり、実際に映画が制作・判断されるプロフェッショナルな映画言語の基準に基づいていないため、映画品質の技巧ではなく、基本的な動画の妥当性を評価している。そこで我々は、映画アカデミー伝統のプロフェッショナルな映画言語に基づき、北京電影学院および華景デジタルメディア&エンターテインメントグループ映画スタジオの監督・教員と共同開発した、テキストから動画へ(T2V)および参照から動画へ(R2V)のベンチマーク「FilmBench」を導入する。このベンチマークは3つの選択に基づいている。第一に、プロンプトはプロの監督が選んだ20の映画ジャンルにわたる受賞作品のクリップからリバースエンジニアリングされており、各プロンプトは検証済みの実写参照映像に紐づけられている。プロンプトは実際のショットリストに従い、大多数が複数ショット(1,169のプロンプト中1,056がマルチショット)であり、従来の単一クリップのベンチマークとは異なる。第二に、評価は3軸、12要素、35(T2V)+3(R2Vのみ)のサブメトリクスからなる3段階の映画分類法に従う。第三に、社内の専門家グレードの自動評価エージェントを開発し、その中核となる映画言語オペレーター群(FilmOps)をオープンソース化する。主要な動画生成モデル(T2V用9、R2V用7)をベンチマークしたところ、評価器はモデルレベルのスピアマンρ = 0.95(T2V)、0.96(R2V)で人間のモデルランキングを再現した。スコアは従来のWebスタイルのベンチマークを大きく下回り、動的美学における一貫したギャップと、単一ショットから複数ショットへの顕著な性能低下が見られ、その低下は弱いモデルほど大きい。
English
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman ho = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.