VGI-BENCH: Probing Visual Intelligence in Video Generation Models

Xuan He, Cong Wei, Yuhao Cheng, Linrui Ma, Yuxuan Zhang, Zuojun Li, Yuhao Wen, Jize Jiang, Zeyi Liu, Yuren Hao, Songcheng Cai, Keming Wu, Penghui Du, Kai Zou, Rui Yang, Chenkai Sun, Ke Yang, Ping Nie, Kelsey R Allen, Chenglong Wang, Michel Galley, Jianfeng Gao, ChengXiang Zhai

arXiv:2608.19583 · 2026-08-27 공개 · arXiv · PDF

video-generation fine-tuning model-evaluation zero-shot visual-reasoning denoising self-correction task-taxonomy

Abstract

Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid evolving processes rather than only plausible final states, and calibrate task difficulty to remain challenging yet partly feasible. To this end, we introduce VGI-bench, containing 27 tasks and 810 instances, organized by a two-level taxonomy of task domains and skill tags for fine-grained evaluation of visual reasoning capabilities of video generation models. Our evaluations show that current generative systems can solve a subset of visually grounded reasoning tasks, but remain far from reliable, with even the strongest model, Seedance~2.0, achieving only 51.0% under our evaluation criteria. Our analysis further explore the output failure modes, input condition sensitivity, performance transfer boundary from synthetic fine-tuning, and internal denoising perspective revealing limited self-correction, where later steps mainly refine early hypotheses rather than correct reasoning errors. We hope VGI-bench will help stimulate the development of next-generation video generation models. We will release our code and data.

한국어 요약

한 줄 요약

VGI-Bench는 영상 생성 모델의 시각적 추론 능력을 평가하기 위한 새로운 벤치마크로, Seedance 2.0 모델이 51.0% 성능을 기록하는 등 현재 모델들의 한계를 드러낸다.

핵심 기여도

핵심 아이디어

VGI-Bench는 기존 영상 생성 모델 평가의 주요 한계를 해결하기 위해 설계되었다. 기존 벤치마크는 추상적 입력이나 최종 상태만을 기준으로 하여, 실제 영상 생성 모델의 시각적 사전 지식과 추론 능력을 정확히 평가하지 못했다. VGI-Bench는 **광학적으로 사실적인 입력**(photorealistic-style inputs)을 사용하고, **중간 과정**(intermediate trajectory)을 기반으로 성공 여부를 판단함으로써, 모델이 단순히 최종 상태를 맞추는 것이 아닌, **시각적 추론을 통해 과정을 정확히 재현하는 능력**을 평가한다.

또한, VGI-Bench는 **2단계 분류 체계**(도메인 + 스킬 태그)를 도입하여, 모델의 특정 능력(예: 물리 법칙 준수, 객체 일관성 유지 등)을 세분화하여 평가할 수 있도록 설계되었다. 이는 기존 벤치마크가 단일 지표로 평가하는 한계를 극복한다.

기술적 접근법

주요 결과

의의 및 한계

VGI-Bench는 영상 생성 모델이 단순한 시각 시뮬레이터를 넘어, **시각적 추론 능력을 갖춘지**를 평가하는 첫 번째 시도로, **다양한 실패 원인을 분석**함으로써 모델 개선 방향을 제시한다. 특히, **입력 조건 민감도**, **미세조정 전이 한계**, **디노이징 동향**을 분석함으로써, 기존 평가 방법의 한계를 극복하고, **과정 중심 평가**(process-sensitive evaluation)를 가능하게 한다는 점에서 학술적 의의가 크다.

하지만, VGI-Bench는 **데이터가 공개되지 않았으며**, 일부 통계는 논문에서 유추한 것이므로, 외부 연구자들의 재현 및 확장이 제한될 수 있다. 또한, **모든 실패가 추론 오류 때문이 아니라 입력 조건에 따른 것일 수 있으므로**, 평가 시 조건 민감도를 보완하는 방향이 필요하다.

실용적 활용

VGI-Bench는 **영상 생성 모델의 시각적 추론 능력을 정량적으로 평가**할 수 있는 도구로, **다음 세대 영상 생성 모델 및 멀티모달 기초 모델**(foundation model)의 개발에 기여할 수 있다. 특히, **시뮬레이션 기반 학습**(embodied learning), **제어 가능한 편집**(controllable editing), **시나리오 기반 생성**(scenario-based generation) 등에서 활용 가능하다.