MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao, Yining Li, Dahua Lin, Kai Chen

arXiv:2406.14515 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation vision-language-models video-understanding temporal-reasoning vlm-evaluation multi-shot videoqa long-form

Abstract

The advent of large vision-language models (LVLMs) has spurred research into their applications in multi-modal contexts, particularly in video understanding. Traditional VideoQA benchmarks, despite providing quantitative metrics, often fail to encompass the full spectrum of video content and inadequately assess models' temporal comprehension. To address these limitations, we introduce MMBench-Video, a quantitative benchmark designed to rigorously evaluate LVLMs' proficiency in video understanding. MMBench-Video incorporates lengthy videos from YouTube and employs free-form questions, mirroring practical use cases. The benchmark is meticulously crafted to probe the models' temporal reasoning skills, with all questions human-annotated according to a carefully constructed ability taxonomy. We employ GPT-4 for automated assessment, demonstrating superior accuracy and robustness over earlier LLM-based evaluations. Utilizing MMBench-Video, we have conducted comprehensive evaluations that include both proprietary and open-source LVLMs for images and videos. MMBench-Video stands as a valuable resource for the research community, facilitating improved evaluation of LVLMs and catalyzing progress in the field of video understanding. The evalutation code of MMBench-Video will be integrated into VLMEvalKit: https://github.com/open-compass/VLMEvalKit.

한국어 요약

한 줄 요약

MMBench-Video는 장시간 영상과 다양한 능력을 기반으로 LVLM의 비디오 이해력을 체계적으로 평가하는 새로운 벤치마크이다.

핵심 기여도

핵심 아이디어

기존 VideoQA 벤치마크는 대부분 짧은 영상(1분 미만)에 초점을 맞추고, 시간적 맥락을 제대로 평가하지 못한다는 한계가 있었다. MMBench-Video는 이 문제를 해결하기 위해 실제 웹 사용자들이 접하는 30초에서 6분까지의 장시간 영상을 사용하며, 시간적 추론 능력을 평가하는 ‘temporal indispensable’ 질문을 포함시켰다. 또한, 기존 GPT-3.5 기반 평가가 인간 판단과의 일치도가 낮다는 문제를 인식하고, GPT-4를 도입하여 평가의 신뢰도를 높였다. 이는 특히 QA 쌍의 길이와 스타일이 다양하기 때문에 의미 있는 평가를 가능하게 한다는 점에서 혁신적이다.

기술적 접근법

주요 결과

의의 및 한계

MMBench-Video는 기존 VideoQA 벤치마크의 한계를 극복하고, 장시간 영상과 다양한 능력을 기반으로 LVLM의 비디오 이해력을 체계적으로 평가할 수 있는 중요한 도구로 기능한다. 특히, GPT-4 기반 평가 시스템은 인간 판단과의 일치도를 높여 평가의 신뢰도를 향상시켰다. 그러나, 일부 영상의 품질이나 질문의 다양성에 대한 한계는 여전히 존재하며, 더 많은 데이터와 평가 방법의 개선이 필요하다. 또한, 일부 프로퍼티어리 모델의 평가 결과는 공개되지 않았다.

실용적 활용

MMBench-Video는 영상 기반 대화형 AI, 영상 요약, 콘텐츠 추천 시스템 등에서 LVLM의 성능을 정확히 평가하는 데 활용될 수 있다. 특히, 시간적 맥락을 강조하는 영상 콘텐츠(예: 뉴스, 교육 영상) 분석에 적합하며, 연구자들이 모델의 시간적 추론 능력을 개선하는 데 중요한 기준이 될 수 있다.