FilmBench: 필름급 시네마틱 비디오 생성을 위한 벤치마크
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
July 27, 2026
저자: Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi, Fei Ding, Weixu Qiao, Jinlin Wang, Xiaotong Lv, Peng Han, Zimeng Li, Fanshu Ding, Yushu Wang, Han Wu, Jingjing Chen, Chongxiao Wang, Yanhao Wu, Chenglong Huang, Xiaoqian Zhu, Jie Tian, Hua Li, Jingjing Fan, Mingshuang Tang, Zhong Li, Hengxia Qiang, Weibin Chen, Jinyang Zhen, Bing Zhao, Lin Qu, Jing Li, Hu Wei
cs.AI
초록
비디오 생성 기술의 발전으로 AI가 생성한 영상과 전문적으로 제작된 영상 간의 시각적 격차가 계속 좁혀지고 있지만, 대부분의 벤치마크는 여전히 웹 소스나 LLM 템플릿에서 프롬프트를 가져오고, 훈련되지 않은 범용 멀티모달 모델로 점수를 매기고 있다. 더 근본적으로, 이러한 평가 분류 체계는 영화가 실제로 제작되고 평가되는 전문적인 시네마틱 언어 기준이 아닌 기본적인 수준(전반적 시각 품질, 대략적인 텍스트 정렬, 시간적 매끄러움)에 머물러 있어, 영화 수준의 장인정신이 아닌 기본적인 비디오 개연성만을 평가한다. 이에 우리는 FilmBench를 소개한다. FilmBench는 영화 아카데미 전통의 전문 시네마틱 언어에 기반한 텍스트-투-비디오(T2V) 및 참조-투-비디오(R2V) 벤치마크로, 베이징 영화 아카데미의 감독 및 교수진, 그리고 호징 디지털 미디어 & 엔터테인먼트 그룹 영화 스튜디오와 공동 개발되었다. 이 벤치마크는 세 가지 선택에 기반한다. 첫째, 프롬프트는 전문 감독이 선정한 20개 영화 장르에 걸친 수상작 영화 클립에서 역설계되었으므로, 모든 프롬프트는 검증된 실사 참조 영상에 기반한다. 프롬프트는 실제 촬영 리스트를 따르며, 대부분이 다중 샷(1,169개 프롬프트 중 1,056개가 다중 샷)으로, 이전의 단일 클립 벤치마크와 차별화된다. 둘째, 평가는 3개 축, 12개 구성 요소, 35개(T2V) + 3개(R2V 전용) 하위 지표로 구성된 3단계 시네마틱 분류 체계를 따른다. 셋째, 우리는 사내 전문가 수준의 자동 평가 에이전트를 개발하고, 그 핵심 도구 모음인 시네마틱 언어 연산자(FilmOps)를 오픈소스로 공개한다. 주요 비디오 생성 모델(T2V 9개, R2V 7개)을 벤치마킹한 결과, 평가자는 모델 수준 스피어만 상관계수 ρ = 0.95(T2V) 및 0.96(R2V)으로 인간 모델 순위를 재현했다. 점수는 기존 웹 스타일 벤치마크보다 훨씬 낮았으며, 동적 미학에서 일관된 격차와 단일 샷에서 다중 샷으로의 현저한 성능 저하(약한 모델일수록 더 두드러짐)가 관찰되었다.
English
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman ho = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.