LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Shaolei Zhang, Qingkai Fang, Zhe Yang, Yang Feng

arXiv:2501.03895 · 2026-07-27 공개 · arXiv · PDF

low-latency token-compression multimodal-model llm-backbone image-understanding flop-reduction video-processing llava-mini

Abstract

The advent of real-time large multimodal models (LMMs) like GPT-4o has sparked considerable interest in efficient LMMs. LMM frameworks typically encode visual inputs into vision tokens (continuous representations) and integrate them and textual instructions into the context of large language models (LLMs), where large-scale parameters and numerous context tokens (predominantly vision tokens) result in substantial computational overhead. Previous efforts towards efficient LMMs always focus on replacing the LLM backbone with smaller models, while neglecting the crucial issue of token quantity. In this paper, we introduce LLaVA-Mini, an efficient LMM with minimal vision tokens. To achieve a high compression ratio of vision tokens while preserving visual information, we first analyze how LMMs understand vision tokens and find that most vision tokens only play a crucial role in the early layers of LLM backbone, where they mainly fuse visual information into text tokens. Building on this finding, LLaVA-Mini introduces modality pre-fusion to fuse visual information into text tokens in advance, thereby facilitating the extreme compression of vision tokens fed to LLM backbone into one token. LLaVA-Mini is a unified large multimodal model that can support the understanding of images, high-resolution images, and videos in an efficient manner. Experiments across 11 image-based and 7 video-based benchmarks demonstrate that LLaVA-Mini outperforms LLaVA-v1.5 with just 1 vision token instead of 576. Efficiency analyses reveal that LLaVA-Mini can reduce FLOPs by 77%, deliver low-latency responses within 40 milliseconds, and process over 10,000 frames of video on the GPU hardware with 24GB of memory.

한국어 요약

한 줄 요약

LLaVA-Mini는 1개의 시각 토큰만으로도 LLaVA-v1.5와 유사한 성능을 달성하는 효율적인 멀티모달 모델이다.

핵심 기여도

핵심 아이디어

LLaVA-Mini는 기존 멀티모달 모델에서 시각 토큰 수가 많아지는 문제를 해결하기 위해, 토큰 수 자체를 극단적으로 줄이는 방식을 제안한다. 기존 연구는 주로 LLM 백본을 작게 만드는 데 집중했으나, LLaVA-Mini는 토큰 수를 줄이는 방향으로 접근한다. 연구자들은 LLM의 초기 레이어에서만 시각 토큰이 중요한 역할을 하며, 이후 레이어에서는 텍스트 토큰이 주요 역할을 한다는 점을 발견했다. 이에 따라, LLaVA-Mini는 텍스트 토큰에 시각 정보를 미리 퓨전하는 **모달리티 프리-퓨전**(modality pre-fusion) 모듈을 도입하여, LLM 백본으로 전달되는 시각 토큰 수를 1개로 줄였다. 이는 토큰 수를 줄이면서도 시각 정보 손실을 최소화하는 핵심 아이디어이다.

기술적 접근법

주요 결과

의의 및 한계

LLaVA-Mini는 토큰 수를 극단적으로 줄이는 방식으로, 기존 멀티모달 모델의 계산 비용 문제를 해결하며 실시간 멀티모달 상호작용을 가능하게 한다. 특히, 고해상도 이미지나 장시간 동영상 처리 시 효율성이 두드러진다. 그러나 압축 과정에서 일부 시각 정보가 손실될 수 있으며, 압축 비율을 너무 높이면 성능 저하가 발생할 수 있다는 한계가 있다. 또한, 압축 모듈과 프리-퓨전 모듈은 특정 아키텍처에 의존적이므로, 다른 LLM과의 호환성 검증이 필요하다.

실용적 활용

LLaVA-Mini는 실시간 이미지/동영상 분석이 필요한 산업, 예를 들어 스마트 모바일 애플리케이션, 자율주행 시스템, 보안 감시 시스템 등에서 유용하게 활용될 수 있다. 또한, GPU 자원이 제한된 환경에서도 빠른 처리가 가능하여, 클라우드 또는 엣지 기반의 멀티모달 서비스 개발에 적합하다.