VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

Xuan He, Dongfu Jiang, Ge Zhang, Max W.F. Ku, Achint Soni, Sherman Siu, Haonan Chen, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang, Quy Duc Do, Yuansheng Ni, Bohan Lyu, Yaswanth Narsupalli, Rongqi "Richard" Fan, Z. Lyu, Yuchen Lin, Wenhu Chen

arXiv:2406.15252 · 2026-07-27 공개 · arXiv · PDF

video-generation rlhf human-feedback video-quality-assessment mantis evalcrafter videoscore automatic-metrics

Abstract

The recent years have witnessed great advances in video generation. However, the development of automatic video metrics is lagging significantly behind. None of the existing metric is able to provide reliable scores over generated videos. The main barrier is the lack of large-scale human-annotated dataset. In this paper, we release VideoFeedback, the first large-scale dataset containing human-provided multi-aspect score over 37.6K synthesized videos from 11 existing video generative models. We train VideoScore (initialized from Mantis)based on VideoFeedback to enable automatic video quality assessment. Experiments show that the Spearman’s correlation betweenVideoScore and humans can reach 77.1 on VideoFeedback-test, beating the prior best metrics by about 50 points. Further result onother held-out EvalCrafter, GenAI-Bench, and VBench show that VideoScore has consistently much higher correlation with humanjudges than other metrics. Due to these results, we believe VideoScore can serve as a great proxy for human raters to (1) rate different video models to track progress (2) simulate fine-grained human feedback in Reinforcement Learning with Human Feedback (RLHF) to improve current video generation models.

한국어 요약

한 줄 요약

VideoScore는 37.6K개의 영상에 대한 인간 평가를 기반으로 학습하여, 기존 지표 대비 50점 이상 높은 상관계수(77.1)를 달성한 자동 영상 평가 모델이다.

핵심 기여도

핵심 아이디어

기존 영상 생성 모델 평가 지표는 단일 측면(예: 시각 품질)에만 초점을 맞추거나, 인간 평가와의 상관성이 낮아 신뢰도가 부족했다. 이에, 본 연구는 **다중 측면 평가**와 **대규모 인간 평가 데이터셋**을 기반으로 한 새로운 자동 평가 모델 VideoScore를 제안한다.

VideoScore는 **Mantis-Idefics2-8B**라는 다중 이미지 및 영상 처리 능력이 뛰어난 모델을 기반으로, VideoFeedback 데이터셋에서 학습된다. 이 데이터셋은 5개의 평가 항목(Visual Quality, Temporal Consistency, Dynamic Degree, Text-to-Video Alignment, Factual Consistency)으로 구성되어 있으며, 각 항목은 1(불만족)에서 4(완벽)까지의 점수를 부여받는다.

기존 MLLM 기반 평가 방식(예: GPT-4o, Gemini)은 인간 평가와의 상관성이 낮았으나, VideoScore는 Mantis 기반 모델을 사용함으로써 **12.1%의 성능 향상**을 달성하는 것으로 나타났다. 이는 모델이 영상의 세부 측면을 정확히 이해하고 평가할 수 있음을 시사한다.

기술적 접근법

주요 결과

의의 및 한계

VideoScore는 기존 영상 평가 지표의 한계를 극복하고, 인간 평가와 높은 상관성을 보이는 자동 평가 모델로, **영상 생성 모델의 성능 추적 및 RLHF에서의 인간 피드백 시뮬레이션**에 유용하게 활용될 수 있다. 특히, 다중 측면 평가와 대규모 인간 평가 데이터셋을 기반으로 한 접근법은 영상 생성 분야의 평가 기준을 혁신적으로 발전시켰다.

그러나, VideoScore는 여전히 **인간 평가와의 완전한 일치**는 이루지 못하며, 평가 항목 간의 **일관성**(IAA: 60% 이상)도 개선 여지가 있다. 또한, 모델이 학습된 데이터셋의 범위를 벗어난 새로운 영상에 대한 일반화 능력도 추가 연구가 필요하다.

실용적 활용

VideoScore는 영상 생성 모델의 성능 비교, 모델 개선을 위한 RLHF 피드백 시뮬레이션, 생성 영상의 질적 평가 등에 활용될 수 있다. 특히, **영상 생성 기업**이나 **AI 연구소**에서 모델 개선 과정에서 자동 평가를 통해 인간 평가 대체 및 효율성 향상에 기여할 수 있다.