LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, K. Saifullah, Sreemanti Dey, Shubh-Agrawal, S. Sandha, Siddartha Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, Micah Goldblum
arXiv:2406.19314 · 2026-07-27 공개 · arXiv · PDF
model-evaluation math-reasoning llm-benchmark open-source-models data-analysis benchmark-updating test-set-contamination livebench
Abstract
Test set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render benchmarks obsolete. To mitigate this, many recent benchmarks crowdsource new prompts and evaluations from human or LLM judges; however, these can introduce significant biases, and break down when scoring hard questions. In this work, we introduce a new benchmark for LLMs designed to be resistant to both test set contamination and the pitfalls of LLM judging and human crowdsourcing. We release LiveBench, the first benchmark that (1) contains frequently-updated questions from recent information sources, (2) scores answers automatically according to objective ground-truth values, and (3) contains a wide variety of challenging tasks, spanning math, coding, reasoning, language, instruction following, and data analysis. To achieve this, LiveBench contains questions that are based on recently-released math competitions, arXiv papers, news articles, and datasets, and it contains harder, contamination-limited versions of tasks from previous benchmarks such as Big-Bench Hard, AMPS, and IFEval. We evaluate many prominent closed-source models, as well as dozens of open-source models ranging from 0.5B to 405B in size. LiveBench is difficult, with top models achieving below 70% accuracy. We release all questions, code, and model answers. Questions are added and updated on a monthly basis, and we release new tasks and harder versions of tasks over time so that LiveBench can distinguish between the capabilities of LLMs as they improve in the future. We welcome community engagement and collaboration for expanding the benchmark tasks and models.
한국어 요약
한 줄 요약
LiveBench는 테스트 세트 오염과 LLM 평가의 편향을 방지하기 위해 설계된, 월별 업데이트와 객관적 평가를 기반으로 한 LLM 벤치마크이다.
핵심 기여도
- LiveBench는 테스트 세트 오염을 방지하기 위해 최근 정보(수학 경시, arXiv, 뉴스 등)에서 유래한 질문을 포함하며, 월별 업데이트한다.
- LLM 또는 인간 평가자의 편향을 제거하기 위해 객관적 정답 기준에 따라 자동 평가한다.
- 기존 벤치마크(Big-Bench Hard, AMPS, IFEval)의 더 어려운 버전을 포함하며, 최상위 모델의 정확도는 70% 미만이다.
- 0.5B에서 405B까지 다양한 크기의 오픈소스 및 클로즈드 모델을 평가하며, 모델 답변, 코드, 질문을 모두 공개한다.
핵심 아이디어
LiveBench는 기존 LLM 벤치마크의 주요 문제인 테스트 세트 오염과 평가 편향을 해결하기 위해 설계되었다. 테스트 세트 오염은 모델이 훈련 시 벤치마크 질문을 이미 학습했을 경우, 성능이 과대평가되는 현상이다. 예를 들어, Codeforces 문제는 훈련 데이터에 포함된 날짜 이후 성능이 급격히 떨어졌으며, GSM8K 수학 문제는 일부 모델이 과적합된 것으로 보인다. LiveBench는 이러한 문제를 해결하기 위해 (1) 월별 업데이트, (2) 객관적 정답 기반 자동 평가, (3) 다양한 어려운 태스크를 포함하는 세 가지 원칙을 따랐다. 특히, 수학, 코딩, 추론, 언어, 명령어 수행, 데이터 분석 등 6개 범주에서 질문을 구성하여 모델의 종합적 능력을 평가한다.
기술적 접근법
- **데이터 출처**: 최근 수학 경시, arXiv 논문, 뉴스, 데이터셋에서 질문을 생성.
- **평가 방식**: LLM 또는 인간 평가자 대신, 객관적 정답에 기반한 자동 평가를 사용.
- **태스크 구성**: Big-Bench Hard, AMPS, IFEval 등 기존 벤치마크의 더 어려운 버전 포함.
- **모델 평가**: 40개 모델(0.5B~405B)을 평가하며, 정확도는 최상위 모델이 70% 미만.
- **업데이트 정책**: 질문은 월별로 업데이트되며, 새로운 태스크와 더 어려운 버전을 순차적으로 추가.
주요 결과
- **모델 성능**: `o1-preview-2024-09-12`가 전체적으로 6% 높은 성능을 보이며 1위.
- **카테고리별 성과**: `o1-preview-2024-09-12`는 데이터 분석, 언어, 수학에서 우수.
- **오픈소스 모델**: `llama-3.1-405b-instruct`와 `qwen2.5-72b-instruct`가 `gpt-4-turbo`를 상회.
- **소형 모델**: `phi-3.5-moe-instruct` (6.6B 파라미터)가 `gpt-3.5`를 초과.
- **월별 업데이트**: 모델 순위 상관관계는 0.997 이상 유지, 평균 점수는 1.2% 감소하며 난이도 상승.
의의 및 한계
LiveBench는 LLM 평가의 공정성과 지속성을 동시에 확보한 첫 번째 벤치마크로, 테스트 세트 오염과 평가 편향 문제를 효과적으로 해결한다. 월별 업데이트와 자동 평가 시스템은 벤치마크의 신뢰도와 유연성을 높인다. 그러나, 모든 질문이 객관적 정답 기반으로 구성되어 있어, 창의성이나 개방형 질문 평가에는 한계가 있다. 또한, 모델 평가 시 자원이 제한되어 50개 이하의 모델만 유지하며, 이는 일부 모델의 평가 누락을 초래할 수 있다.
실용적 활용
LiveBench는 LLM의 지속적 성능 향상을 추적하고, 모델 간 비교를 객관적으로 수행할 수 있는 도구로 활용될 수 있다. 특히, 연구자와 산업계에서 모델 개선 방향을 설정하거나, 새로운 모델의 성능을 검증하는 데 유용하다. 또한, 커뮤니티 기반 업데이트 시스템은 벤치마크의 확장성과 참여도를 높인다.