The advent of real-time large multimodal models (LMMs) like GPT-4o has sparked considerable interest in efficient LMMs. LMM frameworks typically encode visual inputs into vision tokens (continuous representations) and integrate them and textual instructions into the context of large language models (LLMs), where large-scale parameters and numerous context tokens (predominantly vision tokens) result in substantial computational overhead. Previous efforts towards efficient LMMs always focus on replacing the LLM backbone with smaller models, while neglecting the crucial issue of token quantity. In this paper, we introduce LLaVA-Mini, an efficient LMM with minimal vision tokens. To achieve a high compression ratio of vision tokens while preserving visual information, we first analyze how LMMs understand vision tokens and find that most vision tokens only play a crucial role in the early layers of LLM backbone, where they mainly fuse visual information into text tokens. Building on this finding, LLaVA-Mini introduces modality pre-fusion to fuse visual information into text tokens in advance, thereby facilitating the extreme compression of vision tokens fed to LLM backbone into one token. LLaVA-Mini is a unified large multimodal model that can support the understanding of images, high-resolution images, and videos in an efficient manner. Experiments across 11 image-based and 7 video-based benchmarks demonstrate that LLaVA-Mini outperforms LLaVA-v1.5 with just 1 vision token instead of 576. Efficiency analyses reveal that LLaVA-Mini can reduce FLOPs by 77%, deliver low-latency responses within 40 milliseconds, and process over 10,000 frames of video on the GPU hardware with 24GB of memory.
한 줄 요약
LLaVA-Mini는 1개의 시각 토큰만으로도 LLaVA-v1.5와 유사한 성능을 달성하는 효율적인 멀티모달 모델이다.
핵심 기여도
- LLaVA-Mini는 CLIP ViT-L/336px에서 576개의 시각 토큰 대신 1개의 토큰만 사용하면서도 LLaVA-v1.5와 유사한 성능을 보인다.
- 77%의 FLOPs 감소와 40ms 이하의 저지연 응답을 달성한다.
- 24GB GPU 메모리 환경에서 10,000프레임 이상의 동영상을 처리할 수 있다.
- 모달리티 프리-퓨전 모듈과 토큰 압축 모듈을 도입하여 토큰 수를 극단적으로 줄인다.
핵심 아이디어
LLaVA-Mini는 기존 멀티모달 모델에서 시각 토큰 수가 많아지는 문제를 해결하기 위해, 토큰 수 자체를 극단적으로 줄이는 방식을 제안한다. 기존 연구는 주로 LLM 백본을 작게 만드는 데 집중했으나, LLaVA-Mini는 토큰 수를 줄이는 방향으로 접근한다. 연구자들은 LLM의 초기 레이어에서만 시각 토큰이 중요한 역할을 하며, 이후 레이어에서는 텍스트 토큰이 주요 역할을 한다는 점을 발견했다. 이에 따라, LLaVA-Mini는 텍스트 토큰에 시각 정보를 미리 퓨전하는 **모달리티 프리-퓨전**(modality pre-fusion) 모듈을 도입하여, LLM 백본으로 전달되는 시각 토큰 수를 1개로 줄였다. 이는 토큰 수를 줄이면서도 시각 정보 손실을 최소화하는 핵심 아이디어이다.
기술적 접근법
- **모달리티 프리-퓨전 모듈**: 텍스트 토큰에 시각 정보를 미리 퓨전하여, LLM 초기 레이어에서의 퓨전 과정을 대체한다.
- **쿼리 기반 토큰 압축 모듈**: 시각 토큰을 $C \times C$ 크기의 압축 토큰으로 줄이며, 2D sinusoidal positional encoding을 도입해 공간 정보를 보존한다.
- **CLIP ViT-L/336px**를 사용하여 이미지에서 576개의 시각 토큰을 추출하지만, 이를 1개의 토큰으로 압축하여 LLM에 입력한다.
- 텍스트 토큰과 압축된 시각 토큰을 결합해 LLM에 전달하며, 전체 토큰 수는 $C^2 + l_q$로 제한된다.
주요 결과
- 11개의 이미지 기반 벤치마크와 7개의 동영상 기반 벤치마크에서 LLaVA-v1.5 대비 1개의 시각 토큰만으로 유사한 성능을 달성.
- FLOPs 감소율 77%, GPU 메모리 사용량 360MB → 0.6MB (이미지당).
- 이미지 이해 지연 시간 100ms → 40ms 감소.
- 24GB GPU 메모리에서 10,000프레임 이상의 동영상 처리 가능.
의의 및 한계
LLaVA-Mini는 토큰 수를 극단적으로 줄이는 방식으로, 기존 멀티모달 모델의 계산 비용 문제를 해결하며 실시간 멀티모달 상호작용을 가능하게 한다. 특히, 고해상도 이미지나 장시간 동영상 처리 시 효율성이 두드러진다. 그러나 압축 과정에서 일부 시각 정보가 손실될 수 있으며, 압축 비율을 너무 높이면 성능 저하가 발생할 수 있다는 한계가 있다. 또한, 압축 모듈과 프리-퓨전 모듈은 특정 아키텍처에 의존적이므로, 다른 LLM과의 호환성 검증이 필요하다.
실용적 활용
LLaVA-Mini는 실시간 이미지/동영상 분석이 필요한 산업, 예를 들어 스마트 모바일 애플리케이션, 자율주행 시스템, 보안 감시 시스템 등에서 유용하게 활용될 수 있다. 또한, GPU 자원이 제한된 환경에서도 빠른 처리가 가능하여, 클라우드 또는 엣지 기반의 멀티모달 서비스 개발에 적합하다.