OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Z. Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, Maosong Sun

arXiv:2402.14008 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation scientific-benchmark expert-annotations olympiad-bench bilingual-multimodal physics-problems mathematics-problems agii-research

Abstract

Recent advancements have seen Large Language Models (LLMs) and Large Multimodal Models (LMMs) surpassing general human capabilities in various tasks, approaching the proficiency level of human experts across multiple domains. With traditional benchmarks becoming less challenging for these models, new rigorous challenges are essential to gauge their advanced abilities. In this work, we present OlympiadBench, an Olympiad-level bilingual multimodal scientific benchmark, featuring 8,476 problems from Olympiad-level mathematics and physics competitions, including the Chinese college entrance exam. Each problem is detailed with expert-level annotations for step-by-step reasoning. Evaluating top-tier models on OlympiadBench, we implement a comprehensive assessment methodology to accurately evaluate model responses. Notably, the best-performing model, GPT-4V, attains an average score of 17.97% on OlympiadBench, with a mere 10.74% in physics, highlighting the benchmark rigor and the intricacy of physical reasoning. Our analysis orienting GPT-4V points out prevalent issues with hallucinations, knowledge omissions, and logical fallacies. We hope that our challenging benchmark can serve as a valuable resource for helping future AGI research endeavors. The data and evaluation code are available at \url{https://github.com/OpenBMB/OlympiadBench}

한국어 요약

한 줄 요약

OlympiadBench는 수학·물리 올림피아드 수준의 이중 언어·다중 모달 과학 문제를 포함한 8,952개 문제로 구성된 AGI 연구를 위한 도전적인 벤치마크이다.

핵심 기여도

핵심 아이디어

OlympiadBench는 기존 벤치마크가 LLM/LMM에 대해 과도하게 쉬워졌다는 문제를 해결하기 위해 설계되었다. 특히, 수학과 물리의 전문가 수준 문제를 다중 모달 형식으로 구성함으로써, 모델의 과학적 추론 능력을 보다 엄격하게 평가할 수 있도록 했다. 예를 들어, 기하학 문제나 실험 설계 이해는 텍스트만으로는 처리하기 어려운 상황에서 다중 모달 추론이 필수적이다. 본 연구는 GPT-4V와 같은 최고 수준의 LMM이 여전히 계산 오류, 부정확한 추론, 홀로지네이션(hallucination) 등의 문제를 겪는다는 점을 통해, AGI 연구에서의 새로운 도전 과제를 제시한다.

기술적 접근법

주요 결과

의의 및 한계

OlympiadBench는 AGI 연구에서 모델의 과학적 추론 능력을 평가하는 데 중요한 역할을 할 수 있다. 특히, 다중 모달 및 전문가 수준 문제를 포함한 벤치마크는 기존 연구의 한계를 보완할 수 있다. 그러나, 현재 데이터셋은 수학과 물리에만 제한되어 있으며, 다른 과학 분야로의 확장이 필요하다. 또한, 모델 평가 시 자동 평가 파이프라인의 한계가 있을 수 있으므로, 인간 평가와의 비교가 추가 연구 주제가 될 수 있다.

실용적 활용

OlympiadBench는 과학 연구 지원 AI 개발, 교육용 AI 도구 개선, AGI 연구의 성능 평가에 활용될 수 있다. 특히, 수학 및 물리 문제 해결 능력을 향상시키기 위한 모델 개선 연구에 중요한 기초 자료가 될 수 있다.