ChatPaper.aiChatPaper

Sci-VBench:評估科學領域中知識密集與推理密集的影片生成

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

August 10, 2026
作者: Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang, Yilun Zhao
cs.AI

摘要

我們介紹 Sci-VBench,一個全面的基準測試,用於評估跨科學領域中知識密集與推理密集的影片生成。它包含 1,253 個專家註釋的範例,涵蓋四個核心學科下的 60 個主題:自然科學、醫療保健、人文與社會科學,以及工程。每個範例要求模型生成具有豐富時間動態的影片,這些影片需要科學推理和基於知識的綜合,超越表面層次的視覺合理性。我們進一步建立了一個基於評分標準的評估協議。我們的分析顯示,在此協議下,非專家的人工評估者和 MLLM-as-Judge 系統都能與專家判斷達成相對較高的一致性,支持大規模可重現的評估。我們對 16 個前沿的專有與開源模型進行基準測試,發現雖然自動感知品質分數在各系統間緊密聚集,但在提示對齊(Prompt Grounding)和科學與因果正確性(Scientific and Causal Correctness)方面的表現差異甚大,且專有與開源模型之間存在顯著差距。這些發現表明,視覺真實性的進展尚未轉化為對科學與因果動態的可靠建模。
English
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.