LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, K. Saifullah, Sreemanti Dey, Shubh-Agrawal, S. Sandha, Siddartha Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, Micah Goldblum

arXiv:2406.19314 · 2026-07-27 공개 · arXiv · PDF

model-evaluation math-reasoning llm-benchmark open-source-models data-analysis benchmark-updating test-set-contamination livebench

Abstract

Test set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render benchmarks obsolete. To mitigate this, many recent benchmarks crowdsource new prompts and evaluations from human or LLM judges; however, these can introduce significant biases, and break down when scoring hard questions. In this work, we introduce a new benchmark for LLMs designed to be resistant to both test set contamination and the pitfalls of LLM judging and human crowdsourcing. We release LiveBench, the first benchmark that (1) contains frequently-updated questions from recent information sources, (2) scores answers automatically according to objective ground-truth values, and (3) contains a wide variety of challenging tasks, spanning math, coding, reasoning, language, instruction following, and data analysis. To achieve this, LiveBench contains questions that are based on recently-released math competitions, arXiv papers, news articles, and datasets, and it contains harder, contamination-limited versions of tasks from previous benchmarks such as Big-Bench Hard, AMPS, and IFEval. We evaluate many prominent closed-source models, as well as dozens of open-source models ranging from 0.5B to 405B in size. LiveBench is difficult, with top models achieving below 70% accuracy. We release all questions, code, and model answers. Questions are added and updated on a monthly basis, and we release new tasks and harder versions of tasks over time so that LiveBench can distinguish between the capabilities of LLMs as they improve in the future. We welcome community engagement and collaboration for expanding the benchmark tasks and models.

한국어 요약

한 줄 요약

LiveBench는 테스트 세트 오염과 LLM 평가의 편향을 방지하기 위해 설계된, 월별 업데이트와 객관적 평가를 기반으로 한 LLM 벤치마크이다.

핵심 기여도

핵심 아이디어

LiveBench는 기존 LLM 벤치마크의 주요 문제인 테스트 세트 오염과 평가 편향을 해결하기 위해 설계되었다. 테스트 세트 오염은 모델이 훈련 시 벤치마크 질문을 이미 학습했을 경우, 성능이 과대평가되는 현상이다. 예를 들어, Codeforces 문제는 훈련 데이터에 포함된 날짜 이후 성능이 급격히 떨어졌으며, GSM8K 수학 문제는 일부 모델이 과적합된 것으로 보인다. LiveBench는 이러한 문제를 해결하기 위해 (1) 월별 업데이트, (2) 객관적 정답 기반 자동 평가, (3) 다양한 어려운 태스크를 포함하는 세 가지 원칙을 따랐다. 특히, 수학, 코딩, 추론, 언어, 명령어 수행, 데이터 분석 등 6개 범주에서 질문을 구성하여 모델의 종합적 능력을 평가한다.

기술적 접근법

주요 결과

의의 및 한계

LiveBench는 LLM 평가의 공정성과 지속성을 동시에 확보한 첫 번째 벤치마크로, 테스트 세트 오염과 평가 편향 문제를 효과적으로 해결한다. 월별 업데이트와 자동 평가 시스템은 벤치마크의 신뢰도와 유연성을 높인다. 그러나, 모든 질문이 객관적 정답 기반으로 구성되어 있어, 창의성이나 개방형 질문 평가에는 한계가 있다. 또한, 모델 평가 시 자원이 제한되어 50개 이하의 모델만 유지하며, 이는 일부 모델의 평가 누락을 초래할 수 있다.

실용적 활용

LiveBench는 LLM의 지속적 성능 향상을 추적하고, 모델 간 비교를 객관적으로 수행할 수 있는 도구로 활용될 수 있다. 특히, 연구자와 산업계에서 모델 개선 방향을 설정하거나, 새로운 모델의 성능을 검증하는 데 유용하다. 또한, 커뮤니티 기반 업데이트 시스템은 벤치마크의 확장성과 참여도를 높인다.