MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Yuxin Zuo, Shang Qu, Yifei Li, Zhangren Chen, Xuekai Zhu, Ermo Hua, Kaiyan Zhang, Ning Ding, Bowen Zhou

arXiv:2501.18362 · 2026-07-27 공개 · arXiv · PDF

model-evaluation data-synthesis multimodal-evaluation clinical-ai medqa medical-specialties medical-benchmark expert-reasoning

Abstract

We introduce MedXpertQA, a highly challenging and comprehensive benchmark to evaluate expert-level medical knowledge and advanced reasoning. MedXpertQA includes 4,460 questions spanning 17 specialties and 11 body systems. It includes two subsets, Text for text evaluation and MM for multimodal evaluation. Notably, MM introduces expert-level exam questions with diverse images and rich clinical information, including patient records and examination results, setting it apart from traditional medical multimodal benchmarks with simple QA pairs generated from image captions. MedXpertQA applies rigorous filtering and augmentation to address the insufficient difficulty of existing benchmarks like MedQA, and incorporates specialty board questions to improve clinical relevance and comprehensiveness. We perform data synthesis to mitigate data leakage risk and conduct multiple rounds of expert reviews to ensure accuracy and reliability. We evaluate 18 leading models on \benchmark. Moreover, medicine is deeply connected to real-world decision-making, providing a rich and representative setting for assessing reasoning abilities beyond mathematics and code. To this end, we develop a reasoning-oriented subset to facilitate the assessment of o1-like models. Code and data are available at: https://github.com/TsinghuaC3I/MedXpertQA

한국어 요약

한 줄 요약

MedXpertQA는 4,460개의 전문의 수준 질문을 포함한, 의료 분야의 전문 지식과 고급 추론 능력을 평가하는 종합적 벤치마크이다.

핵심 기여도

핵심 아이디어

MedXpertQA는 기존 의료 벤치마크의 한계를 극복하기 위해 설계되었다. 기존 텍스트 기반 벤치마크는 진단 과정의 복잡성을 반영하지 못하며, 다중 모달 벤치마크는 실제 임상 정보를 결합하지 못한 단순 QA 쌍에 머무는 경우가 많다. MedXpertQA는 이에 대응해, 실제 전문의 시험 문제와 임상 기록을 기반으로 한 MM 서브셋을 도입하여, 진단 및 치료 계획 수립에 필요한 전문 지식과 추론 능력을 종합적으로 평가할 수 있도록 설계되었다. 특히, MM 서브셋은 NEJM Image Challenges와 같은 이미지 풍부한 자료를 활용해, 실제 진단 시 겪는 다양한 시각 정보를 시뮬레이션한다.

기술적 접근법

주요 결과

의의 및 한계

MedXpertQA는 의료 AI의 진단 능력과 추론 능력을 종합적으로 평가할 수 있는 체계적인 벤치마크로, 의료 AI의 신뢰성 향상에 기여할 수 있다. 특히, 실제 임상 정보를 반영한 MM 서브셋은 기존의 단순 이미지 기반 QA와 차별화되며, 의료 AI의 실용성 평가에 중요한 역할을 할 수 있다. 그러나 본 연구는 주로 미국 의료 시스템을 기반으로 설계되었기 때문에, 다른 국가의 의료 환경과의 호환성은 명시되지 않았으며, 추가적인 국제적 검증이 필요하다.

실용적 활용

MedXpertQA는 의료 AI 모델의 전문 지식과 추론 능력을 평가하는 데 활용될 수 있으며, 특히 병원, 의료 기관, 연구소에서 의료 AI의 신뢰도와 정확도를 검증하는 데 유용하다. 또한, 의료 전문가 교육 및 시험 문제 개발에도 활용 가능하다.