MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max W.F. Ku, Kai Wang, Alex Zhuang, Rongqi "Richard" Fan, Xiang Yue, Wenhu Chen

arXiv:2406.01574 · 2026-07-27 공개 · arXiv · PDF

large-language-models language-models chain-of-thought model-evaluation reasoning multi-task mmlu-pro prompt-sensitivity

Abstract

In the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in language comprehension and reasoning across diverse domains. However, as models continue to improve, their performance on these benchmarks has begun to plateau, making it increasingly difficult to discern differences in model capabilities. This paper introduces MMLU-Pro, an enhanced dataset designed to extend the mostly knowledge-driven MMLU benchmark by integrating more challenging, reasoning-focused questions and expanding the choice set from four to ten options. Additionally, MMLU-Pro eliminates the trivial and noisy questions in MMLU. Our experimental results show that MMLU-Pro not only raises the challenge, causing a significant drop in accuracy by 16% to 33% compared to MMLU but also demonstrates greater stability under varying prompts. With 24 different prompt styles tested, the sensitivity of model scores to prompt variations decreased from 4-5% in MMLU to just 2% in MMLU-Pro. Additionally, we found that models utilizing Chain of Thought (CoT) reasoning achieved better performance on MMLU-Pro compared to direct answering, which is in stark contrast to the findings on the original MMLU, indicating that MMLU-Pro includes more complex reasoning questions. Our assessments confirm that MMLU-Pro is a more discriminative benchmark to better track progress in the field.

한국어 요약

한 줄 요약

MMLU-Pro는 기존 MMLU 대비 16~33% 낮은 정확도를 기록하며, 추론 능력을 더 잘 평가하는 다분야 언어이해 벤치마크이다.

핵심 기여도

핵심 아이디어

MMLU-Pro는 기존 MMLU 벤치마크가 모델 성능 향상에 따라 포화 상태에 도달했고, 질문의 선택지 수가 적어 모델이 단서를 활용해 정답을 유추할 수 있다는 문제점을 해결하기 위해 설계되었다. MMLU-Pro는 선택지를 4개에서 10개로 확장하고, 더 복잡한 대학 수준 문제를 포함하여 추론 능력을 평가하는 데 초점을 맞추고 있다. 또한, 전문가 검토와 최신 LLM을 활용한 오류 탐지 과정을 거쳐 데이터셋의 신뢰도를 높였다. CoT 추론이 MMLU-Pro에서 성능 향상에 기여하는 반면, MMLU에서는 오히려 성능을 저하시키는 경향이 있어, MMLU-Pro가 추론 기반 평가에 적합함을 보여준다.

기술적 접근법

주요 결과

의의 및 한계

MMLU-Pro는 기존 MMLU가 모델 성능 향상에 따라 포화 상태에 도달한 문제를 해결하고, 추론 능력을 더 잘 평가할 수 있는 새로운 벤치마크로 제시된다. 특히, CoT 추론이 성능 향상에 기여하는 점은 MMLU-Pro가 단순 지식 기반 질문이 아닌 복잡한 추론을 요구함을 보여준다. 또한, 프롬프트 변화에 대한 민감도가 낮아 평가의 안정성을 높였다. 그러나 MMLU-Pro는 여전히 일부 분야에서 모델이 정답을 유추할 수 있는 가능성은 남아 있으며, 더 많은 데이터셋 확장과 오류 검증이 필요하다.

실용적 활용

MMLU-Pro는 대형 언어 모델의 추론 능력을 평가하는 데 적합한 벤치마크로, 연구자들이 모델의 실제 지능 수준을 더 정확히 측정할 수 있도록 지원한다. 특히, 교육, 법률, 공학 등 전문 분야에서 모델의 전문성과 추론 능력을 평가할 때 유용하게 활용될 수 있다.