Many-Shot In-Context Learning
Rishabh Agarwal, Avi Singh, Lei M. Zhang, Bernd Bohnet, Stephanie Chan, Biao Zhang, Ankesh Anand, Zaheer Abbas, Azade Nova, John D. Co-Reyes, Eric Chu, Feryal M. P. Behbahani, Aleksandra Faust, H. Larochelle
arXiv:2404.11018 · 2026-07-27 공개 · arXiv · PDF
llm fine-tuning in-context-learning inference-cost next-token-prediction many-shot reinforced-icl unsupervised-icl
Abstract
Large language models (LLMs) excel at few-shot in-context learning (ICL) -- learning from a few examples provided in context at inference, without any weight updates. Newly expanded context windows allow us to investigate ICL with hundreds or thousands of examples -- the many-shot regime. Going from few-shot to many-shot, we observe significant performance gains across a wide variety of generative and discriminative tasks. While promising, many-shot ICL can be bottlenecked by the available amount of human-generated examples. To mitigate this limitation, we explore two new settings: Reinforced and Unsupervised ICL. Reinforced ICL uses model-generated chain-of-thought rationales in place of human examples. Unsupervised ICL removes rationales from the prompt altogether, and prompts the model only with domain-specific questions. We find that both Reinforced and Unsupervised ICL can be quite effective in the many-shot regime, particularly on complex reasoning tasks. Finally, we demonstrate that, unlike few-shot learning, many-shot learning is effective at overriding pretraining biases, can learn high-dimensional functions with numerical inputs, and performs comparably to fine-tuning. We also find that inference cost increases linearly in the many-shot regime, and frontier LLMs benefit from many-shot ICL to varying degrees. Our analysis also reveals the limitations of next-token prediction loss as an indicator of downstream ICL performance.
한국어 요약
한 줄 요약
Gemini 1.5 Pro를 활용해 수천~수만 샷의 Many-Shot ICL을 실험하고, Reinforced 및 Unsupervised ICL을 제안하여 인간 라벨 의존도를 줄인다.
핵심 기여도
- Gemini 1.5 Pro에서 최대 1M 토큰, 8192 샷의 ICL 성능을 평가.
- Reinforced ICL: 모델 생성 라벨로 인간 라벨 대체.
- Unsupervised ICL: 라벨 없이 문제만 제공.
- Many-Shot ICL이 사전 학습 편향을 극복하고, Fine-tuning 수준 성능 달성.
핵심 아이디어
기존 Few-Shot ICL은 샷 수가 제한적이었으나, 최근 LLM의 컨텍스트 길이가 100배 증가함에 따라 Many-Shot ICL이 가능해졌다. 연구팀은 샷 수가 증가할수록 성능이 급격히 향상되며, 최대 10만 토큰 이상이 필요할 수 있음을 보여준다. 이는 기존 연구에서 고려되지 않았던 ICL의 새로운 가능성을 제시한다. 또한, 인간 라벨 생성 샘플의 부족을 해결하기 위해 Reinforced ICL과 Unsupervised ICL을 제안한다. Reinforced ICL은 모델이 생성한 사고 과정 라벨을 사용하며, Unsupervised ICL은 라벨 없이 문제만 제공함으로써 인간 라벨 의존도를 줄인다.
기술적 접근법
- **모델**: Gemini 1.5 Pro (최대 1M 토큰 컨텍스트)
- **데이터셋**: MATH, GSM8K, XSum, XLSum, GPQA, Big-Bench Hard 등
- **ICL 샷 수**: 최대 8192 샷 (약 10만 토큰)
- **추론 최적화**: KV 캐싱 사용
- **Reinforced ICL**: 모델 생성 라벨을 정답 일치 여부로 필터링
- **Unsupervised ICL**: 라벨 제거, 문제만 제공
- **평가 방법**: 여러 랜덤 시드로 샘플링 후 평균 성능 계산
주요 결과
- MATH, GPQA, Big-Bench Hard 등 복잡한 추론 태스크에서 Reinforced ICL이 Human ICL보다 더 높은 정확도를 보임.
- Many-Shot ICL은 사전 학습 편향을 극복하고, Fine-tuning 수준 성능 달성.
- Next-token prediction loss는 ICL 성능 예측에 신뢰도가 낮음.
- 샷 수가 증가할수록 성능 향상이 지속되며, 최대 10만 토큰 이상이 필요할 수 있음.
의의 및 한계
Many-Shot ICL은 LLM이 사전 학습 데이터와 다른 도메인에도 적응할 가능성을 열며, Fine-tuning 대안으로서의 잠재력을 보여준다. 특히 복잡한 수학 문제 해결과 고차원 함수 학습에서 뛰어난 성능을 보인다. 그러나 인간 라벨 없이도 충분한 성능을 내는 경우가 제한적이며, 모델 생성 라벨의 질에 따라 결과가 크게 달라질 수 있다. 또한, ICL 성능 예측을 위한 loss 지표가 부정확하다는 한계도 드러낸다.
실용적 활용
Many-Shot ICL은 라벨링 비용이 높은 산업 분야 (예: 법률, 의료)에서 유용하게 활용될 수 있으며, 빠르게 변화하는 도메인에서 실시간 적응이 필요한 상황에도 적합하다. 특히, Reinforced ICL은 라벨링 자원이 부족한 저자원 언어나 NLP 이외의 예측 태스크에도 적용 가능하다.