Adaptive Keyframe Sampling for Long Video Understanding

Xi Tang, Jihao Qiu, Lingxi Xie, Yunjie Tian, Jianbin Jiao, Qixiang Ye

arXiv:2502.21271 · 2026-07-27 공개 · arXiv · PDF

video-understanding long-video multimodal-llms video-qa keyframe-selection adaptive-keyframe-sampling information-pre-filtering token-sampling

Abstract

Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. However, when the visual input changes from a single image to a long video, the above paradigm encounters difficulty because the vast amount of video tokens has significantly exceeded the maximal capacity of MLLMs. Therefore, existing video-based MLLMs are mostly established upon sampling a small portion of tokens from input data, which can cause key information to be lost and thus produce incorrect answers. This paper presents a simple yet effective algorithm named Adaptive Keyframe Sampling (AKS). It inserts a plug-and-play module known as keyframe selection, which aims to maximize the useful information with a fixed number of video tokens. We formulate keyframe selection as an optimization involving (1) the relevance between the keyframes and the prompt, and (2) the coverage of the keyframes over the video, and present an adaptive algorithm to approximate the best solution. Experiments on two long video understanding benchmarks validate that AKS improves video QA accuracy (beyond strong baselines) upon selecting informative keyframes. Our study reveals the importance of information pre-filtering in video-based MLLMs. Our codes are available at https://github.com/ncTimTang/AKS

한국어 요약

한 줄 요약

Adaptive Keyframe Sampling(AKS)는 장시간 동영상 이해를 위한 MLLM의 토큰 선택 정확도를 향상시키는 모듈로, LongVideoBench와 VideoMME에서 기존 기준선 대비 높은 정확도를 달성한다.

핵심 기여도

핵심 아이디어

기존 MLLM은 장시간 동영상의 토큰 수가 모델 용량을 초과하므로, 키프레임을 선택적으로 샘플링하여 입력으로 사용한다. 그러나 단순한 균일 샘플링은 중요한 정보를 누락시킬 수 있으며, 이는 오답을 유발한다. 본 연구는 키프레임 선택을 **관련성**과 **커버리지**의 최적화 문제로 정의한다. 관련성은 키프레임과 질문 간의 시각-언어 모델(VL 모델) 기반 유사도로 측정하고, 커버리지는 동영상 내 키프레임의 분포 균형을 재귀적 분할을 통해 추정한다. 이 두 요소를 적절히 균형 잡아 최적의 키프레임 집합을 선택하는 것이 AKS의 핵심 아이디어이다.

기술적 접근법

주요 결과

의의 및 한계

AKS는 MLLM이 고차원 시각 데이터(예: 장시간 동영상)를 효과적으로 처리하기 위해 **정보 사전 필터링**의 중요성을 입증한다. 기존 연구는 키프레임 선택을 간단한 샘플링으로 처리했으나, AKS는 관련성과 커버리지를 동시에 고려함으로써 키프레임의 정보량을 극대화한다. 이는 MLLM의 장기 기억력과 추론 능력을 향상시키는 데 기여할 수 있다.

그러나 AKS는 VL 모델을 기반으로 관련성을 계산하므로, VL 모델의 품질에 따라 성능이 변동될 수 있다. 또한, 재귀적 분할 기반 커버리지 추정은 계산 비용이 증가할 수 있으며, 실시간 처리 환경에서는 한계가 있을 수 있다.

실용적 활용

AKS는 동영상 QA, 요약, 감정 분석 등 장시간 동영상 처리가 필요한 MLLM 기반 시스템에 적용 가능하다. 예를 들어, 온라인 교육 플랫폼에서 동영상 강의의 핵심 내용을 추출하거나, 보안 감시 시스템에서 이상 행동 탐지에 활용할 수 있다. 또한, 4D 데이터(예: 시공간 데이터) 처리에도 AKS의 아이디어가 확장 가능하다.