MLVU: Benchmarking Multi-task Long Video Understanding

Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Shitao Xiao, Xi Yang, Yongping Xiong, Bo Zhang, Tiejun Huang, Zheng Liu

arXiv:2406.04264 · 2026-07-27 공개 · arXiv · PDF

video-understanding long-video context-length egocentric-video llm-backbone image-understanding mlvm multitask

Abstract

The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in video types and evaluation tasks, and the inappropriateness for evaluating LVU performances. To address the above problems, we propose a new benchmark called MLVU (Multitask Long Video Understanding Benchmark) for the comprehensive and in-depth evaluation of LVU. MLVU presents the following critical values: 1) The substantial and flexible extension of video lengths, which enables the benchmark to evaluate LVU performance across a wide range of durations. 2) The inclusion of various video genres, such as movies, surveillance, egocentric videos, and cartoons, reflects the models’ LVU performances in different scenarios. 3) The development of diversified evaluation tasks, which enables a comprehensive examination of MLLMs’ key abilities in long-video understanding. The empirical study with 23 latest MLLMs reveals significant room for improvement in today’s technique, as all existing methods struggle with most of the evaluation tasks and exhibit severe performance degradation when handling longer videos. Additionally, it suggests that factors such as context length, image-understanding ability, and the choice of LLM backbone can play critical roles in future advancements. We anticipate that MLVU will advance the research of LVU by providing a comprehensive and in-depth analysis of MLLMs. The code and dataset can be accessed from https://github.com/JUNJIE99/MLVU.

한국어 요약

한 줄 요약

MLVU는 장시간 동영상 이해 능력을 종합적으로 평가하기 위한 다중 작업 벤치마크로, 기존 한계를 극복하고 MLLM의 성능을 체계적으로 분석한다.

핵심 기여도

핵심 아이디어

MLVU는 기존 동영상 이해 벤치마크가 짧은 동영상, 단일 장르, 단일 작업에 제한되어 있다는 문제를 해결하기 위해 설계되었다. 기존 연구는 대부분 10~15초 이하의 짧은 동영상에 초점을 맞추었고, 평가 작업도 단일 태스크 중심이었다. MLVU는 이에 반해 3분에서 2시간까지 긴 동영상(평균 15분)을 사용하며, 영화, 게임, 일상 영상 등 다양한 장르를 포함한다. 또한, 다중 선택형과 개방형 생성 작업을 포함한 9개의 평가 작업을 통해 MLLM이 전체 동영상의 글로벌 정보와 특정 클립의 로컬 정보를 모두 활용할 수 있는지 평가한다. 이는 MLLM이 실제 장시간 동영상에서 복잡한 정보를 추출하고 종합적으로 이해하는 능력을 측정하는 데 기여한다.

기술적 접근법

MLVU는 다음과 같은 기술적 특징을 갖는다:

주요 결과

의의 및 한계

MLVU는 기존 동영상 이해 벤치마크의 주요 한계를 극복하고, MLLM의 장시간 동영상 이해 능력을 체계적으로 평가할 수 있는 새로운 기준을 제시한다. 특히, 다양한 동영상 길이와 장르, 평가 작업을 통해 MLLM의 실제 적용 가능성과 한계를 명확히 파악할 수 있다. 그러나 MLVU는 아직 초기 단계의 벤치마크이며, 더 많은 연구와 개선이 필요하다. 또한, 평가 작업의 복잡성과 데이터셋의 크기 때문에 일부 모델이 처리에 어려움을 겪을 수 있다.

실용적 활용

MLVU는 영화 분석, 보안 영상 해석, 일상 영상 요약 등 다양한 산업 분야에서 MLLM의 장시간 동영상 이해 능력을 평가하는 데 활용될 수 있다. 또한, 연구자들이 MLLM의 이미지 이해 능력, 컨텍스트 길이, LLM 백본 선택 등 핵심 요소를 개선하는 데 중요한 기초 자료가 될 수 있다.