long-context llm-training mmlu mistral-7b context-utilization narrativeqa information-intensive lost-in-the-middle
Abstract
While many contemporary large language models (LLMs) can process lengthy input, they still struggle to fully utilize information within the long context, known as the lost-in-the-middle challenge. We hypothesize that it stems from insufficient explicit supervision during the long-context training, which fails to emphasize that any position in a long context can hold crucial information. Based on this intuition, our study presents information-intensive (IN2) training, a purely data-driven solution to overcome lost-in-the-middle. Specifically, IN2 training leverages a synthesized long-context question-answer dataset, where the answer requires (1) fine-grained information awareness on a short segment (~128 tokens) within a synthesized long context (4K-32K tokens), and (2) the integration and reasoning of information from two or more short segments. Through applying this information-intensive training on Mistral-7B, we present FILM-7B (FILl-in-the-Middle). To thoroughly assess the ability of FILM-7B for utilizing long contexts, we design three probing tasks that encompass various context styles (document, code, and structured-data context) and information retrieval patterns (forward, backward, and bi-directional retrieval). The probing results demonstrate that FILM-7B can robustly retrieve information from different positions in its 32K context window. Beyond these probing tasks, FILM-7B significantly improves the performance on real-world long-context tasks (e.g., 23.5->26.9 F1 score on NarrativeQA), while maintaining a comparable performance on short-context tasks (e.g., 59.3->59.2 accuracy on MMLU). Github Link: https://github.com/microsoft/FILM.
한국어 요약
한 줄 요약
FILM-7B는 IN2 훈련을 통해 32K 토큰 컨텍스트 내 중간 정보를 효과적으로 활용하는 LLM을 제시한다.
핵심 기여도
- **IN2 훈련**을 제안하여, 4K~32K 토큰 길이의 합성 컨텍스트 내 128 토큰 단위 정보를 강조.
- **FILM-7B** 모델은 32K 토큰 컨텍스트에서 전방, 후방, 양방향 정보 추출 성능 향상.
- **Needle-in-the-Haystack** 등 탐색 실험에서 성능 개선 (예: NarrativeQA F1 23.5 → 26.9).
- **MMLU 정확도 유지** (59.3 → 59.2)로 단축 컨텍스트 성능 저하 없음.
핵심 아이디어
기존 LLM은 훈련 과정에서 컨텍스트 중간 정보를 무시하는 **lost-in-the-middle** 문제가 발생한다. 이는 훈련 데이터에서 위치 편향이 생기기 때문으로, 시작과 끝에 있는 토큰에 더 많은 영향을 받는다. 이를 해결하기 위해, 연구팀은 **IN2 훈련**을 제안한다. 이는 합성된 4K~32K 토큰 길이의 컨텍스트 내에서 **128 토큰 단위의 정보 인식**과 **다중 세그먼트 정보 통합 및 추론**을 요구하는 QA 데이터셋을 사용한다. GPT-4-Turbo를 활용해 QA 쌍을 생성하고, 이를 통해 모델이 컨텍스트 내 어디에 위치하든 정보를 인식하도록 훈련한다.
기술적 접근법
- **IN2 훈련**: 합성된 4K~32K 토큰 길이의 컨텍스트를 128 토큰 단위의 세그먼트로 구성.
- **QA 생성**: GPT-4-Turbo를 사용하여 두 가지 유형의 질문 생성:
1. 단일 세그먼트 내의 **세부 정보 인식** (fine-grained information awareness).
2. 두 개 이상의 세그먼트 정보를 **통합 및 추론**하는 질문.
- **모델**: Mistral-7B 기반으로 훈련하여 **FILM-7B** 모델 생성.
- **추가 실험**: 문서, 코드, 구조화된 데이터 컨텍스트를 포함한 3가지 탐색 패턴 (전방, 후방, 양방향)을 평가.
주요 결과
- **Needle-in-the-Haystack 탐색 실험**: FILM-7B은 32K 토큰 컨텍스트 내 중간 정보를 정확히 추출.
- **실제 세계 태스크 성능**: NarrativeQA에서 F1 점수 23.5 → 26.9 (베이스라인 대비 +3.4).
- **단축 컨텍스트 유지**: MMLU 정확도 59.3 → 59.2 (0.1% 하락, 실질적으로 무시 가능).
- **3가지 탐색 패턴**: 전방, 후방, 양방향 모두에서 성능 향상.
의의 및 한계
FILM-7B은 open-source 모델을 기반으로 하여, **32K 토큰 길이의 컨텍스트 내 정보를 균형 있게 인식**할 수 있는 모델을 제시한다. 이는 기존 LLM의 lost-in-the-middle 문제를 해결하고, **GPT-4-Turbo와 유사한 성능**을 달성함으로써 open-source 모델의 경쟁력을 높인다. 그러나, 합성 데이터에 기반한 훈련이 실제 세계 데이터에 얼마나 잘 일반화되는지는 추가 연구가 필요하다. 또한, 다른 모델 아키텍처에 적용 가능성도 검증이 필요하다.
실용적 활용
FILM-7B은 **문서 분석**, **코드 이해**, **구조화된 데이터 추출** 등 긴 컨텍스트가 필요한 산업 및 연구 분야에 적용 가능하다. 특히, **법률 문서**, **의료 기록**, **소프트웨어 문서** 등에서 정보 추출 성능 향상이 기대된다.