Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Rongrong Ji, Xing Sun
arXiv:2405.21075 · 2026-07-27 공개 · arXiv · PDF
long-context model-evaluation video-mme multi-modal-llms gpt-4o video-analysis subtitle-audio gemini-1-5-pro
Abstract
In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image understanding. The potential of MLLMs to process sequential visual data is still insufficiently explored, highlighting the lack of a comprehensive, high-quality assessment of their performance. In this paper, we introduce Video-MME, the first-ever full-spectrum, Multi-Modal Evaluation benchmark of MLLMs in Video analysis. Our work distinguishes from existing benchmarks through four key features: 1) Diversity in video types, spanning 6 primary visual domains with 30 subfields to ensure broad scenario generalizability; 2) Duration in temporal dimension, encompassing both short-, medium-, and long-term videos, ranging from 11 seconds to 1 hour, for robust contextual dynamics; 3) Breadth in data modalities, integrating multi-modal inputs besides video frames, including subtitles and audios, to unveil the all-round capabilities of MLLMs; 4) Quality in annotations, utilizing rigorous manual labeling by expert annotators to facilitate precise and reliable model assessment. With Video-MME, we extensively evaluate various state-of-the-art MLLMs, and reveal that Gemini 1.5 Pro is the best-performing commercial model, significantly outperforming the open-source models with an average accuracy of 75%, compared to 71.9% for GPT-4o. The results also demonstrate that Video-MME is a universal benchmark that applies to both image and video MLLMs. Further analysis indicates that subtitle and audio information could significantly enhance video understanding. Besides, a decline in MLLM performance is observed as video duration increases for all models. Our dataset along with these findings underscores the need for further improvements in handling longer sequences and multi-modal data, shedding light on future MLLM development. Project page: https://video-m
한국어 요약
한 줄 요약
Video-MME는 비디오 분석에서 MLLM의 다중 모달 성능을 평가하는 최초의 종합적 벤치마크로, 900개의 영상과 2,700개의 QA 쌍으로 구성되어 있다.
핵심 기여도
- Video-MME는 6개 주요 영상 도메인과 30개 하위 분야를 포함한 다종의 영상 유형을 갖춘 최초의 벤치마크.
- 영상 길이가 11초에서 1시간까지 다양하게 구성되어 시간적 맥락을 평가.
- 자막과 오디오를 포함한 다중 모달 입력을 통합하여 MLLM의 전반적 능력을 평가.
- 전문가가 수작업으로 라벨링한 2,700개의 QA 쌍을 포함하여 정확도 높은 평가 가능.
핵심 아이디어
기존 MLLM 평가가 정적 이미지에 집중한 반면, Video-MME는 동적 비디오 데이터를 기반으로 MLLM의 종합적 성능을 평가하는 데 초점을 맞춘다. 이는 비디오의 시간적 맥락, 다양한 모달 정보(자막, 오디오), 그리고 다양한 도메인을 고려한 평가를 가능하게 한다. 특히, Gemini 1.5 Pro가 75.7%의 정확도를 기록한 반면, LLaVA-NeXT-Video는 52.5%에 그친다는 점에서, 상업적 모델이 개방형 모델보다 비디오 이해력에서 우수함을 보여준다. 또한, 자막과 오디오 정보는 Gemini 1.5 Pro의 정확도를 각각 5.9%와 4.7% 향상시키는 것으로 나타나, 비디오 분석에서의 다중 모달 정보 통합의 중요성을 강조한다.
기술적 접근법
- **데이터셋 구성**: 900개의 영상, 2,700개의 QA 쌍, 평균 길이 1024초.
- **모델 평가**: GPT-4o, Gemini 1.5 Pro, InternVL-Chat-V1.5, LLaVA-NeXT-Video 등 다양한 MLLM 평가.
- **모달 통합**: 영상 프레임 외에도 자막과 오디오 정보를 포함.
- **하이퍼파라미터**: QA 토큰 수 35.6, 자막 토큰 수 3,066.5.
- **분석 방법**: 영상 길이, 모달 정보 유무에 따른 성능 변화를 비교 분석.
주요 결과
- **Gemini 1.5 Pro**: 평균 정확도 75.7%로 최고 성능.
- **LLaVA-NeXT-Video**: 52.5%로 상대적으로 낮은 성능.
- **자막 통합**: Gemini 1.5 Pro의 정확도 5.9% 향상.
- **오디오 통합**: 4.7% 향상.
- **영상 길이 증가**: 모든 모델에서 성능 하락 관찰됨.
- **Video-MME-S/M/L**: 각각 80.8초, 520.2초, 2471.0초의 평균 길이를 가짐.
의의 및 한계
Video-MME는 비디오 분석에서 MLLM의 다중 모달 처리 능력을 체계적으로 평가할 수 있는 최초의 벤치마크로, 비디오 길이와 모달 정보의 영향을 정확히 측정할 수 있다. 특히, 상업적 모델과 개방형 모델 간의 성능 격차를 명확히 보여주며, 자막과 오디오 정보의 통합이 비디오 이해도를 향상시킨다는 점에서 실용적 가치가 있다. 그러나, 영상 길이가 늘어날수록 성능이 감소하는 한계가 있으며, 이는 MLLM이 긴 시퀀스를 처리하는 능력이 부족하다는 점을 시사한다. 또한, 데이터셋 크기(900개 영상)는 대규모 평가에 비해 상대적으로 작을 수 있다.
실용적 활용
Video-MME는 비디오 기반 콘텐츠 분석, 자막 및 오디오 정보 활용을 요구하는 멀티모달 시스템 개발, 그리고 MLLM의 시간적 맥락 이해 능력 향상을 위한 연구에 활용될 수 있다. 특히, 영상 길이 증가에 따른 성능 저하를 개선하기 위한 모델 아키텍처 개발에 중요한 기초 자료가 될 수 있다.