Make Your LLM Fully Utilize the Context

Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng, Jian-Guang Lou

arXiv:2404.16811 · 2026-07-27 공개 · arXiv · PDF

long-context llm-training mmlu mistral-7b context-utilization narrativeqa information-intensive lost-in-the-middle

Abstract

While many contemporary large language models (LLMs) can process lengthy input, they still struggle to fully utilize information within the long context, known as the lost-in-the-middle challenge. We hypothesize that it stems from insufficient explicit supervision during the long-context training, which fails to emphasize that any position in a long context can hold crucial information. Based on this intuition, our study presents information-intensive (IN2) training, a purely data-driven solution to overcome lost-in-the-middle. Specifically, IN2 training leverages a synthesized long-context question-answer dataset, where the answer requires (1) fine-grained information awareness on a short segment (~128 tokens) within a synthesized long context (4K-32K tokens), and (2) the integration and reasoning of information from two or more short segments. Through applying this information-intensive training on Mistral-7B, we present FILM-7B (FILl-in-the-Middle). To thoroughly assess the ability of FILM-7B for utilizing long contexts, we design three probing tasks that encompass various context styles (document, code, and structured-data context) and information retrieval patterns (forward, backward, and bi-directional retrieval). The probing results demonstrate that FILM-7B can robustly retrieve information from different positions in its 32K context window. Beyond these probing tasks, FILM-7B significantly improves the performance on real-world long-context tasks (e.g., 23.5->26.9 F1 score on NarrativeQA), while maintaining a comparable performance on short-context tasks (e.g., 59.3->59.2 accuracy on MMLU). Github Link: https://github.com/microsoft/FILM.

한국어 요약

한 줄 요약

FILM-7B는 IN2 훈련을 통해 32K 토큰 컨텍스트 내 중간 정보를 효과적으로 활용하는 LLM을 제시한다.

핵심 기여도

핵심 아이디어

기존 LLM은 훈련 과정에서 컨텍스트 중간 정보를 무시하는 **lost-in-the-middle** 문제가 발생한다. 이는 훈련 데이터에서 위치 편향이 생기기 때문으로, 시작과 끝에 있는 토큰에 더 많은 영향을 받는다. 이를 해결하기 위해, 연구팀은 **IN2 훈련**을 제안한다. 이는 합성된 4K~32K 토큰 길이의 컨텍스트 내에서 **128 토큰 단위의 정보 인식**과 **다중 세그먼트 정보 통합 및 추론**을 요구하는 QA 데이터셋을 사용한다. GPT-4-Turbo를 활용해 QA 쌍을 생성하고, 이를 통해 모델이 컨텍스트 내 어디에 위치하든 정보를 인식하도록 훈련한다.

기술적 접근법

주요 결과

의의 및 한계

FILM-7B은 open-source 모델을 기반으로 하여, **32K 토큰 길이의 컨텍스트 내 정보를 균형 있게 인식**할 수 있는 모델을 제시한다. 이는 기존 LLM의 lost-in-the-middle 문제를 해결하고, **GPT-4-Turbo와 유사한 성능**을 달성함으로써 open-source 모델의 경쟁력을 높인다. 그러나, 합성 데이터에 기반한 훈련이 실제 세계 데이터에 얼마나 잘 일반화되는지는 추가 연구가 필요하다. 또한, 다른 모델 아키텍처에 적용 가능성도 검증이 필요하다.

실용적 활용

FILM-7B은 **문서 분석**, **코드 이해**, **구조화된 데이터 추출** 등 긴 컨텍스트가 필요한 산업 및 연구 분야에 적용 가능하다. 특히, **법률 문서**, **의료 기록**, **소프트웨어 문서** 등에서 정보 추출 성능 향상이 기대된다.