long-context retrieval-augmented question-answering model-agnostic llm-performance ruler-benchmark coding-tasks prompting-strategy
Abstract
Large language models (LLMs) often fail to scale their performance on long-context tasks performance in line with the context lengths they support. This gap is commonly attributed to retrieval failures -- the models'inability to identify relevant information in the long inputs. Accordingly, recent efforts often focus on evaluating and improving LLMs'retrieval performance: if retrieval is perfect, a model should, in principle, perform just as well on a long input as it does on a short one -- or should it? This paper presents findings that the answer to this question may be negative. Our systematic experiments across 5 open- and closed-source LLMs on math, question answering, and coding tasks reveal that, even when models can perfectly retrieve all relevant information, their performance still degrades substantially (13.9%--85%) as input length increases but remains well within the models'claimed lengths. This failure occurs even when the irrelevant tokens are replaced with minimally distracting whitespace, and, more surprisingly, when they are all masked and the models are forced to attend only to the relevant tokens. A similar performance drop is observed when all relevant evidence is placed immediately before the question. Our findings reveal a previously-unrealized limitation: the sheer length of the input alone can hurt LLM performance, independent of retrieval quality and without any distraction. They motivate our simple, model-agnostic mitigation strategy that transforms a long-context task into a short-context one by prompting the model to recite the retrieved evidence before attempting to solve the problem. On RULER, we observe a consistent improvement of GPT-4o up to 4% on an already strong baseline.
한국어 요약
한 줄 요약
긴 입력 길이 자체가 LLM의 성능을 저하시키며, 이는 단순한 검색 실패와 무관하다.
핵심 기여도
- 5개의 오픈/클로즈드 소스 LLM에서 실험한 결과, 입력 길이가 늘어날수록 성능이 13.9%–85% 감소함.
- 불필요한 토큰을 공백 또는 마스킹 처리해도 성능 저하가 발생함.
- GPT-4o에서 RULER 데이터셋에서 기존 기준 대비 최대 4% 성능 향상.
- 단순한 "retrieve-then-reason" 전략을 제안하여 긴 문맥 작업을 짧은 문맥 작업으로 변환.
핵심 아이디어
기존 연구는 LLM이 긴 문맥 작업에서 실패하는 원인을 주로 "검색 실패"로 보았으나, 본 연구는 **입력 길이 자체가 성능 저하를 유발할 수 있음을 밝힘**.
LLM이 모든 관련 정보를 완벽히 검색(100% 정확도)해도, 입력 길이가 늘어날수록 정확도가 감소하는 현상이 관찰됨.
이를 증명하기 위해, 불필요한 토큰을 **공백** 또는 **마스킹**하여 모델이 관련 정보에만 집중하도록 했음에도 불구하고, 성능 저하가 지속됨.
이러한 실험은 **"입력 길이 자체가 LLM의 추론 능력을 저하시킨다"는 새로운 통찰**을 제시하며, 기존의 "검색-이용" 이분법적 접근법의 한계를 드러냄.
기술적 접근법
- **5개의 LLM**(Llama-3.1-8B Instruct, Mistral-v0.3-7B Instruct 등)을 사용한 실험.
- **MMLU, GSM8K, HumanEval, Var Sum, RULER** 등 다양한 작업에서 성능 평가.
- **"Perfect retrieval" 조건**에서 실험: 관련 정보를 100% 정확히 검색한 상태에서 입력 길이를 증가.
- **불필요한 토큰 처리**: 공백 대체, 마스킹, 관련 정보를 질문 바로 앞에 배치.
- **Mitigation Strategy**: 모델에게 관련 정보를 먼저 "재생"하도록 유도한 후, 질문에 답변하게 함 (retrieve-then-reason).
- **GPT-4o**를 RULER 데이터셋에서 실험하여, 기존 기준 대비 최대 4% 성능 향상.
주요 결과
- **MMLU**에서 Llama-3.1-8B Instruct의 정확도가 24.2% 감소.
- **HumanEval**에서 Mistral-v0.3-7B Instruct의 정확도가 44% 감소.
- **RULER QA1**에서 GPT-4o의 정확도가 기존 88.2–90.4%에서 92.2%로 향상 (4K 토큰 기준).
- **RULER QA2**에서 GPT-4o의 정확도가 최대 4% 향상 (32K 토큰 기준).
- **Var Sum**에서 Llama의 정확도가 96%에서 37%로, Mistral은 68%에서 24%로 감소.
- 입력 길이가 7K 토큰에서 이미 성능 저하가 발생하며, 이는 검색 성능 저하 이전임.
의의 및 한계
- **의의**: 기존의 "검색 실패" 이외에도, 입력 길이 자체가 LLM의 추론 능력을 저하시킨다는 새로운 한계를 밝힘.
- **실용적 가치**: "retrieve-then-reason" 전략은 모델에 무관하며, 단순하고 효과적인 해결책.
- **한계**: 본 연구는 특정 조건(예: 마스킹, 공백 대체)에서 실험했으며, 실제 세계의 복잡한 입력 환경에서는 다른 결과가 나올 수 있음.
- **추가 연구 필요**: 다양한 모델 아키텍처와 입력 유형에서의 일반화 가능성 검증 필요.
실용적 활용
- **대화형 챗봇**에서 긴 대화 기록을 처리할 때, 입력 길이가 성능에 영향을 줄 수 있으므로, **입력 길이를 줄이는 전략**이 유용.
- **코드 생성, QA, 수학 문제 해결** 등에서, **retrieve-then-reason** 전략을 적용하여 성능을 안정적으로 유지할 수 있음.
- **RAG 시스템**에서, 단순히 더 많은 문서를 추가하는 것보다, **입력 길이를 제어하는 전략**이 더 효과적일 수 있음.