DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models

Saeed Ranjbar Alvar, Gursimran Singh, Mohammad Akbari, Yong Zhang

arXiv:2503.02175 · 2026-07-27 공개 · arXiv · PDF

large-language-models multimodal-models latency-reduction token-selection video-language token-pruning visual-tokens divprune

Abstract

Large Multimodal Models (LMMs) have emerged as powerful models capable of understanding various data modalities, including text, images, and videos. LMMs encode both text and visual data into tokens that are then combined and processed by an integrated Large Language Model (LLM). Including visual tokens substantially increases the total token count, often by thousands. The increased input length for LLM significantly raises the complexity of inference, resulting in high latency in LMMs. To address this issue, token pruning methods, which remove part of the visual tokens, are proposed. The existing token pruning methods either require extensive calibration and fine-tuning or rely on suboptimal importance metrics which results in increased redundancy among the retained tokens. In this paper, we first formulate token pruning as Max-Min Diversity Problem (MMDP) where the goal is to select a subset such that the diversity among the selected tokens is maximized. Then, we solve the MMDP to obtain the selected subset and prune the rest. The proposed method, DivPrune, reduces redundancy and achieves the highest diversity of the selected tokens. By ensuring high diversity, the selected tokens better represent the original tokens, enabling effective performance even at high pruning ratios without requiring fine-tuning. Extensive experiments with various LMMs show that DivPrune achieves state-of-the-art accuracy over 16 image- and video-language datasets. Additionally, DivPrune reduces both the end-to-end latency and GPU memory usage for the tested models. The code is available here⋄.

한국어 요약

한 줄 요약

DivPrune는 시각 토큰의 다양성 최대화를 통해 LMM의 추론 속도와 메모리 사용량을 줄이는 토큰 프루닝 방법이다.

핵심 기여도

핵심 아이디어

기존 토큰 프루닝 방법은 주로 어텐션 점수나 중요도 기반으로 토큰을 선택하지만, 이는 토큰 간 중복을 유발하여 성능 저하를 초래한다. DivPrune은 이러한 문제를 해결하기 위해 토큰 간 거리 기반의 Max-Min Diversity Problem(MMDP)을 도입하여, 선택된 토큰 집합의 다양성을 최대화한다. 이는 토큰 간 최소 거리를 증가시키는 방식으로, 선택된 토큰이 원본 토큰을 더 잘 대표하도록 보장한다. 본 연구는 MMDP를 토큰 프루닝에 적용한 최초의 시도로, 기존의 Min-Max 전략이나 랜덤 프루닝과 비교해 15.8% 이상의 성능 향상을 보인다.

기술적 접근법

주요 결과

의의 및 한계

DivPrune은 기존의 중요도 기반 프루닝 방법의 한계를 극복하고, 토큰 간 중복을 줄이며 성능을 유지하는 새로운 접근법을 제시한다. 특히, 모델 미세조정 없이도 높은 프루닝 비율에서도 효과적인 성능을 보이는 점에서 실용적 가치가 높다. 또한, 다양한 LLM 아키텍처와 비전 인코더에 적용 가능하며, 추론 최적화 기법과의 호환성도 뛰어나다. 그러나 본 연구는 특정 LMM 아키텍처에만 적용되었으며, 더 다양한 모델 구조에서의 일반화 가능성은 추가 연구가 필요하다.

실용적 활용

DivPrune는 저연산 비용으로 높은 성능을 유지하는 LMM의 추론 최적화에 적합하다. 특히, 실시간 이미지/비디오 분석이 필요한 IoT, 모바일, 클라우드 환경에서 유용하게 활용될 수 있다. 또한, 모델 미세조정 없이 즉시 적용 가능한 plug-and-play 속성 덕분에, 다양한 산업 분야에서 즉각적인 성능 향상과 리소스 절감을 기대할 수 있다.