Sci-VBench: 科学分野における知識・推論集約型動画生成の評価
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
August 10, 2026
著者: Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang, Yilun Zhao
cs.AI
要旨
本稿では、科学領域における知識集約的かつ推論集約的な動画生成を評価するための包括的ベンチマークであるSci-VBenchを紹介する。本ベンチマークは、自然科学、ヘルスケア、人文・社会科学、工学の4つの中核分野にわたる60の主題を網羅する1,253件の専門家注釈付き事例で構成される。各事例は、表面レベルの視覚的妥当性を超え、科学的推論と知識に基づく合成を要求する、時間的にリッチな動画の生成をモデルに課す。さらに、我々はルーブリックに基づく評価プロトコルを確立する。分析の結果、このプロトコルの下では、非専門家の人間評価者とMLLM-as-Judgeシステムの双方が専門家の判定と比較的高い一致を示し、大規模な再現可能な評価を支援できることが明らかになった。我々は、プロプライエタリおよびオープンソースの最先端モデル16種をベンチマークし、自動知覚品質スコアはシステム間で緊密にクラスタリングされる一方、プロンプト追従性と科学的・因果的正確性は大きく変動し、顕著なプロプライエタリ・オープンソース間格差が見られることを発見した。これらの知見は、視覚的リアリズムの進歩が、科学的・因果的ダイナミクスの信頼性の高いモデリングには未だ結びついていないことを示している。
English
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.