Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu, Baobao Chang

arXiv:2410.07985 · 2026-07-27 공개 · arXiv · PDF

large-language-models model-evaluation mathematical-reasoning llm-benchmark problem-solving human-annotation olympiad-math difficulty-levels

Abstract

Recent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8\% on MATH dataset), indicating their inadequacy for truly challenging these models. To bridge this gap, we propose a comprehensive and challenging benchmark specifically designed to assess LLMs' mathematical reasoning at the Olympiad level. Unlike existing Olympiad-related benchmarks, our dataset focuses exclusively on mathematics and comprises a vast collection of 4428 competition-level problems with rigorous human annotation. These problems are meticulously categorized into over 33 sub-domains and span more than 10 distinct difficulty levels, enabling a holistic assessment of model performance in Olympiad-mathematical reasoning. Furthermore, we conducted an in-depth analysis based on this benchmark. Our experimental results show that even the most advanced models, OpenAI o1-mini and OpenAI o1-preview, struggle with highly challenging Olympiad-level problems, with 60.54\% and 52.55\% accuracy, highlighting significant challenges in Olympiad-level mathematical reasoning.

한국어 요약

한 줄 요약

Omni-MATH는 4,428개의 올림피아드 수준 문제로 구성된 LLM 수학 추론 평가 벤치마크로, 기존 모델이 60.54% 이하 정확도를 보임.

핵심 기여도

핵심 아이디어

기존 수학 벤치마크(GSM8K, MATH)는 LLM이 높은 정확도(예: OpenAI o1이 MATH에서 94.8%)를 보이면서 더 이상 모델 평가에 적합하지 않다. 이에 따라, **올림피아드 수준의 문제**를 중심으로 한 **전면적 수학 추론 평가 시스템**이 필요하다는 통찰이 제시된다. Omni-MATH는 실제 수학 올림피아드의 문제 선정 과정을 반영하여, **문제의 난이도와 분야별 분류**를 통해 LLM의 수학적 추론 능력을 **다차원적으로 평가**할 수 있도록 설계되었다. 특히, **33개의 하위 분야**와 **10개 이상의 난이도 수준**은 기존 벤치마크보다 훨씬 세분화된 평가를 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

Omni-MATH는 기존 수학 벤치마크가 LLM의 수학 추론 능력을 충분히 평가하지 못하는 문제를 해결하고, **올림피아드 수준의 문제 해결 능력**을 측정하는 새로운 기준을 제시한다. 특히, **GPT-4o와 Omni-Judge를 활용한 평가 시스템**은 정확도와 신뢰성을 동시에 확보한 점에서 학술적·실용적 가치가 있다. 그러나, **모든 문제가 텍스트 기반**이며, **다중 모달성**을 고려한 평가가 제한된 점은 한계로 작용할 수 있다. 또한, **Best-of-N 방식의 무효성**은 추후 테스트 타임 스케일링 기법 연구의 필요성을 강조한다.

실용적 활용

Omni-MATH는 수학 올림피아드 문제 해결 능력을 평가하고자 하는 **교육 기관**, **AI 연구소**, **수학 커뮤니티**에서 활용 가능하다. 또한, **LLM의 수학 추론 능력 개선**을 위한 연구 및 모델 최적화에 중요한 기준이 될 수 있다. 특히, **GPT-4o와 Omni-Judge를 활용한 평가 시스템**은 저비용으로도 신뢰도 높은 평가를 가능하게 하므로, **공개 평가 플랫폼**에 적합하다.