LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu, Jun Chen, Chenchen Zhu, Zechun Liu, Fanyi Xiao, Bala Varadarajan, Florian Bordes, Zhuang Liu, Hu Xu, Hyunwoo J. Kim, Bilge Soran, Raghuraman Krishnamoorthi, Mohamed Elhoseiny, Vikas Chandra

arXiv:2410.17434 · 2026-07-27 공개 · arXiv · PDF

video-understanding long-video dino-v2 token-reduction video-language mlvu spatiotemporal-compression cross-modal-query

Abstract

Multimodal Large Language Models (MLLMs) have shown promising progress in understanding and analyzing video content. However, processing long videos remains a significant challenge constrained by LLM's context size. To address this limitation, we propose LongVU, a spatiotemporal adaptive compression mechanism thats reduces the number of video tokens while preserving visual details of long videos. Our idea is based on leveraging cross-modal query and inter-frame dependencies to adaptively reduce temporal and spatial redundancy in videos. Specifically, we leverage DINOv2 features to remove redundant frames that exhibit high similarity. Then we utilize text-guided cross-modal query for selective frame feature reduction. Further, we perform spatial token reduction across frames based on their temporal dependencies. Our adaptive compression strategy effectively processes a large number of frames with little visual information loss within given context length. Our LongVU consistently surpass existing methods across a variety of video understanding benchmarks, especially on hour-long video understanding tasks such as VideoMME and MLVU. Given a light-weight LLM, our LongVU also scales effectively into a smaller size with state-of-the-art video understanding performance.

한국어 요약

한 줄 요약

LongVU는 장시간 동영상 이해를 위해 시공간 적응형 압축 기법을 도입한 MLLM 모델로, DINOv2와 텍스트 유도 쿼리를 활용해 토큰 수를 줄이며 정확도를 유지한다.

핵심 기여도

핵심 아이디어

LongVU는 장시간 동영상의 처리를 위해 **시공간 적응형 압축**을 제안한다. 기존 연구는 고정된 프레임 샘플링을 사용해 동영상의 비균일한 내용(예: 정적 vs 동적 장면)을 무시했으나, LongVU는 **DINOv2**의 시각적 특징을 활용해 유사한 프레임을 제거함으로써 **시계열 중복**을 줄인다. 또한, 텍스트 쿼리에 따라 중요한 프레임은 **전체 토큰(144개)**을 유지하고, 나머지는 **공간 풀링**을 통해 **64개 토큰**으로 축소한다. 이는 시각 정보 손실을 최소화하면서도 컨텍스트 길이 내에서 더 많은 프레임을 처리할 수 있도록 한다. 마지막으로, **시간 의존성 기반 공간 토큰 축소(STC)**를 통해 토큰 수를 동적으로 조절한다.

기술적 접근법

주요 결과

의의 및 한계

LongVU는 기존 MLLM이 처리하지 못한 **1시간 길이 동영상**을 8k 컨텍스트 내에서 처리할 수 있게 하며, **시공간 적응형 압축**을 통해 토큰 수를 줄이면서도 시각 정보 손실을 최소화하는 데 성공했다. 특히, **DINOv2**의 시각적 특징이 **SigLIP**보다 효과적임을 실험적으로 입증했다. 그러나, **16k 컨텍스트**를 지원하는 모델과 비교하면 일부 성능 저하가 발생하며, **복잡한 장면**에서는 압축 과정에서 정보 손실이 발생할 수 있다.

실용적 활용

LongVU는 **장시간 동영상 분석**, **비디오 QA 시스템**, **교육용 콘텐츠 이해** 등에 적용 가능하다. 특히, **가벼운 LLM 기반 모델**로도 뛰어난 성능을 보여주어 **자원 제한 환경**에서 유용하게 활용될 수 있다.