🤖 AI Summary
This study addresses the absence of systematic evaluation benchmarks for knowledge- and reasoning-intensive scientific video generation. The authors introduce the first comprehensive benchmark spanning four scientific domains, comprising 1,253 expert-annotated samples, and propose a scalable rubric-based evaluation protocol that synergistically combines human experts with multimodal large language models (MLLM-as-Judge) to assess model performance across dimensions such as scientific correctness, causal reasoning, and prompt alignment. Evaluation of 16 state-of-the-art models reveals that, despite comparable perceptual quality, they exhibit substantial disparities in scientific reasoning capabilities, with proprietary models significantly outperforming open-source counterparts. These findings highlight a fundamental gap between visual realism and accurate modeling of scientific dynamics in current approaches.
📝 Abstract
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.