Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang, Yilun Zhao

arXiv:2608.09873 · 2026-08-11 공개 · arXiv · PDF

video-generation llm-evaluation benchmarking multimodal-models scientific-reasoning science-education sci-vbench knowledge-grounded

Abstract

We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.

한국어 요약

한 줄 요약

Sci-VBench는 과학 분야에서 지식과 추론이 필요한 동영상 생성 모델을 평가하는 벤치마크로, 1,253개의 전문가 주석이 담긴 예제를 포함한다.

핵심 기여도

핵심 아이디어

기존 동영상 생성 모델 평가가 시각적 사실성에 집중하는 반면, Sci-VBench는 과학적 메커니즘과 인과 관계를 정확히 반영하는 능력을 평가하는 데 초점을 맞춘다. 이는 단순히 시각적으로 설득력 있는 동영상이 아니라, 과학적 제약과 시간적 일관성을 유지하는 동영상 생성을 요구한다. Sci-VBench는 각 예제에 대해 전문가가 작성한 'Reference Guide'와 'Evaluation Rubric'을 제공하여, 비전문가나 MLLM도 동일한 기준으로 평가할 수 있도록 설계되었다. 특히, Scientific and Causal Correctness 차원은 기존의 일반적 품질 평가 기준이 포착하지 못하는 메커니즘 수준의 오류를 측정한다.

기술적 접근법

주요 결과

의의 및 한계

Sci-VBench는 과학 분야에서의 동영상 생성 모델 평가를 체계화하고, 전문가 평가 기준을 비전문가와 MLLM이 재현할 수 있도록 설계된 점에서 혁신적이다. 특히, Prompt Grounding과 Scientific and Causal Correctness 차원에서의 성능 차이는 시각적 사실성과 과학적 정확성 간의 괴리를 드러내며, 모델 개선 방향을 제시한다. 그러나 일부 과학 분야에서는 예제 수가 제한적이고, 모델이 시간적 일관성을 유지하는 데 어려움을 겪는다는 한계가 있다.

실용적 활용

Sci-VBench는 과학 교육, 연구 시뮬레이션, 의료 시각화 등에서 모델의 과학적 정확성과 시간적 일관성을 평가하는 데 활용될 수 있다. 특히, 모델 개발자들이 과학적 메커니즘을 정확히 반영하는 능력을 향상시키는 데 중요한 기준이 될 수 있다.