ChatPaper.aiChatPaper

Sci-VBench: 과학 분야에서 지식 및 추론 집약적 비디오 생성 평가

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

August 10, 2026
저자: Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang, Yilun Zhao
cs.AI

초록

우리는 과학적 영역 전반에 걸쳐 지식 및 추론 집약적 비디오 생성을 평가하기 위한 포괄적인 벤치마크인 Sci-VBench를 소개한다. Sci-VBench는 자연과학, 의료, 인문·사회과학, 공학의 네 가지 핵심 학문 분야에 걸친 60개 주제를 아우르는 1,253개의 전문가 주석 예시로 구성된다. 각 예시는 모델이 표면적 수준의 시각적 그럴듯함을 넘어 과학적 추론과 지식 기반 합성을 요구하는 시간적으로 풍부한 비디오를 생성해야 한다. 또한 우리는 루브릭 기반 평가 프로토콜을 구축한다. 우리의 분석에 따르면, 이 프로토콜 하에서 비전문가 인간 평가자와 MLLM-as-Judge 시스템 모두 전문가 판단과 비교적 높은 일치도를 달성할 수 있어 대규모 재현 가능한 평가를 지원한다. 우리는 16개의 최첨단 독점 및 오픈소스 모델을 벤치마킹한 결과, 자동 지각 품질 점수는 시스템 전반에 걸쳐 밀집하게 분포하는 반면, 프롬프트 충실성과 과학적 및 인과적 정확성에서는 상당한 차이를 보이며 독점-오픈소스 간 격차가 두드러짐을 발견했다. 이러한 결과는 시각적 사실성의 발전이 아직 과학적 및 인과적 역학의 신뢰할 수 있는 모델링으로 이어지지 않았음을 보여준다.
English
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.