MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Ming Yin, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, Graham Neubig

arXiv:2409.02813 · 2026-07-27 공개 · arXiv · PDF

benchmarking chain-of-thought multimodal-benchmark multimodal-ai reasoning-evaluation visual-textual-integration vision-only ocr-prompts

Abstract

This paper introduces MMMU-Pro, a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark. MMMU-Pro rigorously assesses multimodal models' true understanding and reasoning capabilities through a three-step process based on MMMU: (1) filtering out questions answerable by text-only models, (2) augmenting candidate options, and (3) introducing a vision-only input setting where questions are embedded within images. This setting challenges AI to truly"see"and"read"simultaneously, testing a fundamental human cognitive skill of seamlessly integrating visual and textual information. Results show that model performance is substantially lower on MMMU-Pro than on MMMU, ranging from 16.8% to 26.9% across models. We explore the impact of OCR prompts and Chain of Thought (CoT) reasoning, finding that OCR prompts have minimal effect while CoT generally improves performance. MMMU-Pro provides a more rigorous evaluation tool, closely mimicking real-world scenarios and offering valuable directions for future research in multimodal AI.

한국어 요약

한 줄 요약

MMMU-Pro는 MMMU 벤치마크를 기반으로, 시각-언어 통합 능력을 엄격히 평가하는 새로운 멀티모달 평가 기준이다.

핵심 기여도

핵심 아이디어

MMMU-Pro는 기존 MMMU 벤치마크가 단순히 통계적 패턴을 학습한 모델을 과도하게 평가할 수 있다는 문제를 해결하기 위해 설계되었다. 특히, 모델이 단순히 텍스트만으로 답할 수 있는 질문을 필터링하고, 선택지를 4개에서 10개로 증가시켜 추측을 어렵게 만든다. 가장 중요한 혁신은 **시각 입력 설정**(vision-only input setting)을 도입한 점이다. 이 설정에서는 질문이 이미지 내에 포함되어 있어 모델이 동시에 "보기"와 "읽기"를 수행해야 하며, 이는 인간의 핵심 인지 능력인 시각-언어 통합을 시뮬레이션한다. 이는 과학 다이어그램 해석, GUI 탐색 등 실제 세계의 복잡한 시나리오를 반영한다.

기술적 접근법

주요 결과

의의 및 한계

MMMU-Pro는 기존 MMMU가 단순 패턴 학습에 의존하는 모델을 과도하게 평가하는 문제를 해결하며, 실제 세계 시나리오에 더 가까운 평가를 가능하게 한다. 특히, 시각-텍스트 통합 능력을 시험하는 설정은 인간 인지 능력에 근접한 평가 기준을 제시한다. 그러나 MMMU-Pro는 기존 MMMU의 질문을 기반으로 하므로, 새로운 질문 생성 능력을 평가하지는 못한다. 또한, OCR 기술이 발전함에 따라 OCR 프롬프트의 효과가 줄어들고 있어, 단순 텍스트 인식 이상의 능력을 평가하는 새로운 지표 개발이 필요하다.

실용적 활용

MMMU-Pro는 과학, 교육, UI/UX 분야에서 시각-텍스트 통합 능력을 필요로 하는 AI 시스템 개발에 유용한 평가 도구로 활용될 수 있다. 특히, 복잡한 다이어그램 해석, 그래픽 사용자 인터페이스 내 질문 처리, 실제 환경에서의 스크린샷 기반 질의 응답 등에 적용 가능하다.