InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory

Chaojun Xiao, Pengle Zhang, Xu Han, Guangxuan Xiao, Yankai Lin, Zhengyan Zhang, Zhiyuan Liu, Song Han, Maosong Sun

arXiv:2402.04617 · 2026-07-27 공개 · arXiv · PDF

llm long-context training-free attention-mechanism context-window extrapolation memory-based sequence-processing

Abstract

Large language models (LLMs) have emerged as a cornerstone in real-world applications with lengthy streaming inputs (e.g., LLM-driven agents). However, existing LLMs, pre-trained on sequences with a restricted maximum length, cannot process longer sequences due to the out-of-domain and distraction issues. Common solutions often involve continual pre-training on longer sequences, which will introduce expensive computational overhead and uncontrollable change in model capabilities. In this paper, we unveil the intrinsic capacity of LLMs for understanding extremely long sequences without any fine-tuning. To this end, we introduce a training-free memory-based method, InfLLM. Specifically, InfLLM stores distant contexts into additional memory units and employs an efficient mechanism to lookup token-relevant units for attention computation. Thereby, InfLLM allows LLMs to efficiently process long sequences with a limited context window and well capture long-distance dependencies. Without any training, InfLLM enables LLMs that are pre-trained on sequences consisting of a few thousand tokens to achieve comparable performance with competitive baselines that continually train these LLMs on long sequences. Even when the sequence length is scaled to $1,024$K, InfLLM still effectively captures long-distance dependencies. Our code can be found in \url{https://github.com/thunlp/InfLLM}.

한국어 요약

한 줄 요약

InfLLM은 추가 학습 없이 기존 LLM이 긴 시퀀스를 처리할 수 있도록 돕는 메모리 기반의 훈련 없는 방법이다.

핵심 기여도

핵심 아이디어

InfLLM은 기존 LLM이 긴 시퀀스를 처리할 때 발생하는 out-of-domain 및 distraction 문제를 해결하기 위해 훈련 없이 외부 메모리 기반 접근법을 제안한다. 기존 LLM은 훈련 시퀀스 길이가 짧아 긴 입력을 처리하지 못하는 한계가 있었으며, 이를 극복하기 위해 추가 훈련이 필요했으나, 이는 비용이 많이 들고 모델 성능을 불안정하게 만들 수 있었다. InfLLM은 슬라이딩 윈도우 어텐션과 병합된 외부 메모리 모듈을 도입하여, 각 토큰이 관련된 정보만 선택적으로 참조하도록 한다. 이는 어텐션 매트릭스의 희소성을 고려해 불필요한 노이즈를 제거하고, 긴거리 의존성을 효과적으로 포착하는 데 기여한다. 특히, 블록 수준 메모리 단위를 사용해 토큰별 계산을 줄이고, 빈번히 사용되지 않는 단위는 CPU 메모리에 저장함으로써 GPU 메모리 사용량을 줄인다.

기술적 접근법

주요 결과

의의 및 한계

InfLLM은 기존 LLM의 길이 제한을 극복하면서 추가 학습 없이도 긴 시퀀스를 처리할 수 있는 기술적 가능성을 보여준다. 특히, 블록 단위 메모리와 동적 메모리 로드 기법은 계산 효율성과 정확도를 동시에 향상시키며, 실용적인 스트리밍 애플리케이션에 유용할 수 있다. 그러나 InfLLM은 외부 메모리에 의존하기 때문에, 메모리 크기나 접근 속도가 성능에 영향을 줄 수 있다. 또한, 현재는 훈련 없이 적용 가능한 모델에만 제한되며, 더 복잡한 시퀀스 구조나 다중 언어 환경에서의 성능은 추가 실험 필요.

실용적 활용

InfLLM은 LLM 기반 에이전트, 스트리밍 데이터 처리, 대규모 문서 분석 등 긴 시퀀스를 다루는 산업 및 연구 분야에 적용 가능하다. 특히, 추가 학습 없이 기존 모델을 활용할 수 있어, 빠른 개발 주기와 비용 절감이 필요한 상황에서 유리하다.