llm-evaluation benchmarking reasoning-evaluation aime-2024 imo-2025 model-contamination math-arena proof-writing
Abstract
The rapid advancement of reasoning capabilities in large language models (LLMs) has led to notable improvements on mathematical benchmarks. However, many of the most commonly used evaluation datasets (e.g., AIME 2024) are widely available online, making it difficult to disentangle genuine reasoning from potential memorization. Furthermore, these benchmarks do not evaluate proof-writing capabilities, which are crucial for many mathematical tasks. To address this, we introduce MathArena, a new benchmark based on the following key insight: recurring math competitions provide a stream of high-quality, challenging problems that can be used for real-time evaluation of LLMs. By evaluating models as soon as new problems are released, we effectively eliminate the risk of contamination. Using this framework, we find strong signs of contamination in AIME 2024. Nonetheless, evaluations on harder competitions, such as CMIMC 2025, demonstrate impressive reasoning capabilities in top-performing models. MathArena is also the first benchmark for proof-writing capabilities. On IMO 2025, top models achieve slightly less than 40%, demonstrating both notable progress and significant room for improvement. So far, we have evaluated over $50$ models across seven competitions, totaling $162$ problems. As an evolving benchmark, MathArena will continue to track the progress of LLMs on newly released competitions, ensuring rigorous and up-to-date evaluation of mathematical reasoning.
한국어 요약
한 줄 요약
MathArena는 수학 경시 문제를 활용한 LLM 수학 추론 능력 평가 벤치마크로, 오염되지 않은 문제와 증명 작성 평가를 제공한다.
핵심 기여도
- AIME 2024 문제 집합에 강한 오염 징후를 발견함.
- CMIMC 2025, SMT 2025 등 최신 경시 문제를 사용해 LLM의 추론 능력을 평가함.
- USAMO 2025에서 최상위 모델이 25% 미만 성능을 보이며, 증명 작성 평가의 필요성을 강조함.
- 30개 모델을 5개 경시 대회, 149개 문제로 평가함.
핵심 아이디어
기존 수학 평가 데이터셋은 온라인에 공개되어 있어 LLM이 문제를 학습했을 가능성이 높아 진정한 추론 능력을 평가하기 어렵다. MathArena는 매년 개최되는 수학 경시 대회 문제를 실시간으로 평가함으로써 데이터 오염을 최소화하고, 특히 증명 작성 능력을 평가하는 데 초점을 맞춘다. 이는 기존 평가가 단순 최종 답만을 중점적으로 다루는 한계를 보완한다. 예를 들어, USAMO 2025 문제는 LLM이 단순 계산이 아닌 논리적 증명을 요구하므로, 모델의 수학적 깊이를 정확히 평가할 수 있다.
기술적 접근법
- **MathArena 프레임워크**: 경시 문제를 실시간으로 수집하고, 모델의 응답을 자동으로 추출 및 검증하는 파이프라인을 구축함.
- **평가 대상 모델**: o4-mini, Gemini-2.5-Pro, Grok 3, DeepSeek-R1, Qwen3 등 최상위 성능을 내는 모델을 선정함.
- **문제 출처**: AIME 2024, CMIMC 2025, SMT 2025, USAMO 2025 등 2025년에 개최된 경시 대회 문제를 사용함.
- **평가 기준**: 최종 답 정확도와 증명 작성 능력을 병행 평가함.
주요 결과
- **AIME 2024**: 대부분의 LLM이 문제를 학습했을 가능성이 높아, 오염된 평가임.
- **SMT 2025**: 최상위 모델이 높은 추론 능력을 보임.
- **USAMO 2025**: 최상위 모델이 25% 미만 성능을 기록함.
- **총 평가**: 30개 모델, 5개 대회, 149개 문제로 평가됨.
의의 및 한계
MathArena는 수학적 추론과 증명 작성 능력을 동시에 평가하는 첫 번째 벤치마크로, LLM의 진정한 수학 능력을 평가하는 데 기여한다. 특히, 실시간 문제 평가를 통해 데이터 오염을 방지하는 점에서 혁신적이다. 그러나, 일부 경시 문제는 공개되지 않거나, 평가 과정에서 인간의 주관이 개입될 수 있는 한계가 있다. 또한, 증명 작성 평가는 자동 평가가 어려워, 인간 평가자와의 협업이 필요하다.
실용적 활용
MathArena는 수학 교육, 연구, AI 모델 개발 분야에서 LLM의 수학 능력을 객관적으로 평가하는 데 활용될 수 있다. 특히, 모델의 추론 능력과 증명 작성 능력을 동시에 평가할 수 있어, 학문적 연구 및 산업적 응용에서의 신뢰도를 높일 수 있다.