LVBench: An Extreme Long Video Understanding Benchmark

Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, Xiaohan Zhang, Ji Qi, Shiyu Huang, Bin Xu, Yuxiao Dong, Ming Ding, Jie Tang

arXiv:2406.08035 · 2026-07-27 공개 · arXiv · PDF

video-understanding multimodal-models long-term-memory long-video embodied-intelligence information-extraction movie-reviews sports-commentary

Abstract

Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of meeting the demands of real-world applications such as embodied intelligence for long-term decision-making, in-depth movie reviews and discussions, and live sports commentary, all of which require comprehension of long videos spanning several hours. To address this gap, we introduce LVBench, a benchmark specifically designed for long video understanding. Our dataset comprises publicly sourced videos and encompasses a diverse set of tasks aimed at long video comprehension and information extraction. LVBench is designed to challenge multimodal models to demonstrate long-term memory and extended comprehension capabilities. Our extensive evaluations reveal that current multimodal models still underperform on these demanding long video understanding tasks. Through LVBench, we aim to spur the development of more advanced models capable of tackling the complexities of long video comprehension.

한국어 요약

한 줄 요약

LVBench는 4시간 이상 긴 영상 이해 능력을 평가하는 새로운 벤치마크로, 기존 모델의 한계를 드러내며 연구 촉진을 목표로 한다.

핵심 기여도

핵심 아이디어

기존의 멀티모달 대형 언어 모델은 1분 미만의 짧은 영상 이해에 성공했으나, 몇 시간에 이르는 긴 영상에 대한 처리 능력은 여전히 부족하다. 이는 실생활 응용(예: 실시간 스포츠 해설, 영화 분석 등)에서 필수적인 능력이므로, 이를 평가할 수 있는 새로운 벤치마크가 필요하다는 통찰에서 출발했다. LVBench는 6개의 핵심 능력을 기반으로 복합 질문을 구성하여 모델의 장기 기억 및 확장 이해 능력을 체계적으로 평가한다.

기술적 접근법

주요 결과

의의 및 한계

LVBench는 긴 영상 이해 분야에서 체계적인 평가 기준을 제공하며, 모델의 장기 기억 및 복합 정보 추출 능력을 측정하는 데 기여한다. 특히, 데이터 수집 및 어노테이션의 어려움을 극복한 점에서 학술적 가치가 높다. 그러나, 인간 성능에 대한 구체적인 수치가 제시되지 않았으며, 모델 개선 방향에 대한 구체적 제안은 제시되지 않았다는 한계가 있다.

실용적 활용

LVBench는 영화 분석, 스포츠 해설, 장기 의사결정 시스템 등에서 필요한 긴 영상 이해 능력을 평가하는 데 활용될 수 있다. 또한, 대형 멀티모달 모델의 개선을 위한 연구 및 개발에 중요한 기준이 될 수 있다.