OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM

Yutao Hu, Tian-Xin Li, Quanfeng Lu, Wenqi Shao, Junjun He, Yu Qiao, Ping Luo

arXiv:2402.09181 · 2026-07-27 공개 · arXiv · PDF

model-evaluation vision-language-models medical-imaging multimodal-ai medical-vqa real-world-medical-applications large-scale-benchmark medical-datasets

Abstract

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in various multimodal tasks. However, their potential in the medical domain re-mains largely unexplored. A significant challenge arises from the scarcity of diverse medical images spanning various modalities and anatomical regions, which is essential in real-world medical applications. To solve this problem, in this paper, we introduce OmniMedVQA, a novel comprehensive medical Visual Question Answering (VQA) benchmark. This benchmark is collected from 73 different medical datasets, including 12 different modalities and covering more than 20 distinct anatomical regions. Importantly, all images in this benchmark are sourced from authentic medical scenarios, ensuring alignment with the requirements of the medical field and suitability for evaluating LVLMs. Through our extensive experiments, we have found that existing LVLMs struggle to address these medical VQA problems effectively. Moreover, what surprises us is that medical-specialized LVLMs even exhibit inferior performance to those general-domain models, calling for a more versatile and robust LVLM in the biomedical field. The evaluation results not only reveal the current limitations of LVLM in understanding real medical images but also highlight our dataset's significance. Our code with dataset are available at https://github.com/OpenGVLab/ Multi Modality-Arena.

한국어 요약

한 줄 요약

OmniMedVQA는 73개 의료 데이터셋을 기반으로 구성된 대규모 의료 VQA 벤치마크로, LVLM의 의료 분야 성능 평가에 기여한다.

핵심 기여도

핵심 아이디어

기존 LVLM이 의료 분야에서의 평가가 부족하고, 특히 다양한 모달리티와 해부학적 부위를 아우르는 데이터셋이 부족한 점을 문제로 삼았다. OmniMedVQA는 73개 의료 데이터셋을 기반으로 12개의 모달리티(MRI, CT, X-Ray 등)와 20개 이상의 해부학적 부위를 포함한 QA 형식으로 전환하여, LVLM의 의료 이미지 이해 능력을 종합적으로 평가할 수 있는 벤치마크를 제시한다. 특히, GPT의 강력한 추론 능력을 활용해 분류 데이터를 VQA 형식으로 전환함으로써 데이터 수집 비용을 줄이고, 실제 의료 시나리오와 밀접하게 연계된 데이터를 구축했다.

기술적 접근법

주요 결과

의의 및 한계

OmniMedVQA는 의료 분야에서 LVLM의 성능을 종합적으로 평가할 수 있는 첫 번째 대규모 VQA 벤치마크로, 의료 이미지 이해 능력의 한계를 드러내며, 의료 전용 LVLM의 개선 필요성을 강조한다. 특히, CT, MRI 등 일부 모달리티에서는 의료 전용 모델이 우수하지만, 다른 모달리티에서는 일반 모델 못지 않다는 점은 LVLM이 다양한 의료 이미지에 대한 일반화 능력을 갖추어야 한다는 점을 시사한다. 한계로는 일부 의료 전용 모델이 다중 선택형 평가 형식에 적응하지 못하는 문제가 있으며, 이는 평가 형식에 따른 성능 차이를 반영할 수 있다.

실용적 활용

OmniMedVQA는 의료 분야에서 LVLM의 성능 평가 및 모델 개선에 활용될 수 있으며, 특히 의료 이미지 분석, 진단 지원 시스템, 의료 교육 등에서 모델의 신뢰도를 높이는 데 기여할 수 있다. 또한, 다양한 모달리티와 해부학적 부위를 아우르는 데이터셋은 의료 AI 연구자들이 모델의 일반화 능력을 평가하는 데 유용한 기준이 될 수 있다.