LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning

Hongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang, Zirui Liu, Chia-yuan Chang, Huiyuan Chen, Xia Hu

arXiv:2401.01325 · 2026-07-27 공개 · arXiv · PDF

llm long-context attention-mechanism context-window code-modification inference self-extend grouped-attention

Abstract

It is well known that LLMs cannot generalize well to long contexts whose lengths are larger than the training sequence length. This poses challenges when employing LLMs for processing long input sequences during inference. In this work, we argue that LLMs themselves have inherent capabilities to handle long contexts without fine-tuning. To achieve this goal, we propose SelfExtend to extend the context window of LLMs by constructing bi-level attention information: the grouped attention and the neighbor attention. The grouped attention captures the dependencies among tokens that are far apart, while neighbor attention captures dependencies among adjacent tokens within a specified range. The two-level attentions are computed based on the original model's self-attention mechanism during inference. With minor code modification, our SelfExtend can effortlessly extend existing LLMs' context window without any fine-tuning. We conduct comprehensive experiments on multiple benchmarks and the results show that our SelfExtend can effectively extend existing LLMs' context window length. The code can be found at \url{https://github.com/datamllab/LongLM}.

한국어 요약

한 줄 요약

SelfExtend는 미세조정 없이 기존 LLM의 컨텍스트 윈도우를 확장하는 새로운 방법으로, 그룹 어텐션과 이웃 어텐션을 기반으로 한다.

핵심 기여도

핵심 아이디어

기존 LLM은 학습 시 고정된 길이의 시퀀스만 처리하도록 훈련받았기 때문에, 추론 시 긴 입력을 받으면 성능이 급격히 저하된다. 이는 위치 인코딩(Positional Encoding)의 Out-of-Distribution (O.O.D.) 문제로 인해 발생한다. 특히, RoPE 기반 모델에서는 상대 위치 $ m - n $이 학습 시 경험되지 않은 값일 경우 문제가 발생한다.

SelfExtend는 이 문제를 해결하기 위해, 추론 시 발생하는 큰 상대 위치를 학습 시 경험한 값으로 매핑한다. 이는 단순한 `floor` 연산을 통해 이루어지며, 정확한 위치보다는 의미와 순서가 더 중요하다는 언어의 특성을 반영한다. 이 접근법은 T5와 iRPE의 아이디어와 유사하며, 기존 어텐션 메커니즘을 그대로 활용하므로 추가적인 훈련 없이도 가능하다.

기술적 접근법

주요 결과

의의 및 한계

SelfExtend는 기존 LLM의 내재적 능력을 활용하여 컨텍스트 윈도우를 확장하는 새로운 접근법으로, 미세조정 없이도 긴 텍스트 처리가 가능하다는 점에서 학술적·실용적 가치가 있다. 특히, 기존 어텐션 메커니즘을 그대로 활용하므로 추가적인 복잡성 없이 적용 가능하다는 장점이 있다.

그러나, SelfExtend는 기본 구현 시 계산 비용이 증가하며, 그룹 크기가 커질수록 성능이 저하된다. 또한, 전체 시퀀스를 처리하기 때문에 일부 압축 기법(예: 프롬프트 압축)에 비해 계산량이 더 많다. 긴 컨텍스트 평가 방법론도 아직 표준화되지 않아 실험 결과 해석에 어려움이 있을 수 있다.

실용적 활용

SelfExtend는 문서 요약, 법률 분석, 대규모 텍스트 분석 등 긴 입력을 처리해야 하는 산업 및 연구 분야에 적용 가능하다. 특히, 기존 LLM을 최소한의 수정으로 확장할 수 있어, 빠른 개발 주기와 자원 효율성이 필요한 상황에서 유용하다.