Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Yuan Feng, Junlin Lv, Yukun Cao, Xike Xie, S. K. Zhou

arXiv:2407.11550 · 2026-07-27 공개 · arXiv · PDF

kv-cache llm-inference attention-heads long-sequence cache-eviction budget-allocation ruler-dataset longbench-dataset

Abstract

Large Language Models have excelled in various domains but face efficiency challenges due to the growing Key-Value (KV) cache required for long-sequence inference. Recent efforts aim to reduce KV cache size by evicting vast non-critical cache elements during runtime while preserving generation quality. However, these methods typically allocate compression budgets uniformly across all attention heads, ignoring the unique attention patterns of each head. In this paper, we establish a theoretical loss upper bound between pre- and post-eviction attention output, explaining the optimization target of prior cache eviction methods, while guiding the optimization of adaptive budget allocation. Base on this, we propose {\it Ada-KV}, the first head-wise adaptive budget allocation strategy. It offers plug-and-play benefits, enabling seamless integration with prior cache eviction methods. Extensive evaluations on 13 datasets from Ruler and 16 datasets from LongBench, all conducted under both question-aware and question-agnostic scenarios, demonstrate substantial quality improvements over existing methods. Our code is available at https://github.com/FFY0/AdaKV.

한국어 요약

한 줄 요약

Ada-KV는 주의력 헤드별로 적응적 예산을 할당하여 KV 캐시 제거 성능을 향상시키는 최초의 전략이다.

핵심 기여도

핵심 아이디어

기존 KV 캐시 제거 방법은 모든 주의력 헤드에 동일한 Top-k 기반 예산을 할당하지만, 헤드별로 주의력 분포가 다르므로 이는 비효율적이다. 일부 헤드는 주의력이 특정 토큰에 집중되어 있고, 다른 헤드는 토큰 전체에 걸쳐 분산되어 있다. 이에 따라, 희소한 주의력 분포를 가진 헤드는 더 많은 캐시 공간을 필요로 하며, 기존의 균일한 예산 할당은 이 불균형을 무시한다.

Ada-KV는 이러한 헤드별 주의력 패턴을 고려하여, 헤드별로 예산을 조정하는 적응적 할당 전략을 제안한다. 이는 eviction loss 상한식을 최소화하는 방식으로, 헤드별로 Top-k 기준을 다르게 적용하여, 전체적인 제거 손실을 줄인다. 이론적으로, L1 거리 기반의 eviction loss 상한식을 도출하고, 이를 최소화하는 방향으로 예산을 재할당한다.

기술적 접근법

주요 결과

의의 및 한계

Ada-KV는 기존 KV 캐시 제거 방법의 균일한 예산 할당 문제를 해결하고, 헤드별 주의력 패턴을 고려한 적응적 예산 할당을 제안함으로써, 캐시 제거 손실을 줄이고 생성 품질을 향상시킨다. 이는 특히 question-agnostic 시나리오에서 더 큰 효과를 보이며, 실제 적용 시 유용한 전략이다.

그러나, Ada-KV는 헤드별 주의력 패턴을 사전에 분석해야 하므로, 실시간 적용 시 추가 계산 비용이 발생할 수 있다. 또한, 모든 헤드에 동일한 기준을 적용하는 기존 방법 대비 복잡성이 증가하므로, 특정 애플리케이션에서는 성능 향상보다 복잡도 증가가 더 큰 문제가 될 수 있다.

실용적 활용

Ada-KV는 대규모 언어 모델의 효율적인 추론을 필요로 하는 산업 분야, 예를 들어, 대화형 AI, 문서 요약, 코드 생성 등에서 적용 가능하다. 특히, 캐시 메모리가 제한된 GPU 환경에서 실시간 성능 향상이 필요한 경우, SnapKV나 Pyramid와의 통합을 통해 즉시 활용할 수 있다. NVIDIA와 Cloudflare의 실제 프로젝트에서도 이미 적용되고 있으며, 향후 다양한 추론 엔진에 확장 가능하다.