MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Yifan Zhang, Huanyu Zhang, Haochen Tian, Chaoyou Fu, Shuangqing Zhang, Jun Wu, Feng Li, Kun Wang, Qingsong Wen, Zhang Zhang, Liang Wang, Rong Jin, Tien-Ping Tan

arXiv:2408.13257 · 2026-07-27 공개 · arXiv · PDF

mllm multimodal-llm high-resolution llm-benchmark data-collection annotation real-world-scenarios evaluation

Abstract

Comprehensive evaluation of Multimodal Large Language Models (MLLMs) has recently garnered widespread attention in the research community. However, we observe that existing benchmarks present several common barriers that make it difficult to measure the significant challenges that models face in the real world, including: 1) small data scale leads to a large performance variance; 2) reliance on model-based annotations results in restricted data quality; 3) insufficient task difficulty, especially caused by the limited image resolution. To tackle these issues, we introduce MME-RealWorld. Specifically, we collect more than $300$K images from public datasets and the Internet, filtering $13,366$ high-quality images for annotation. This involves the efforts of professional $25$ annotators and $7$ experts in MLLMs, contributing to $29,429$ question-answer pairs that cover $43$ subtasks across $5$ real-world scenarios, extremely challenging even for humans. As far as we know, MME-RealWorld is the largest manually annotated benchmark to date, featuring the highest resolution and a targeted focus on real-world applications. We further conduct a thorough evaluation involving $28$ prominent MLLMs, such as GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet. Our results show that even the most advanced models struggle with our benchmarks, where none of them reach $60\%$ accuracy. The challenges of perceiving high-resolution images and understanding complex real-world scenarios remain urgent issues to be addressed. The data and evaluation code are released at https://mme-realworld.github.io/ .

한국어 요약

한 줄 요약

MME-RealWorld는 13,366개의 고해상도 이미지로 구성된 최대 규모의 수작업 라벨 벤치마크로, 28개 MLLM이 60% 미만의 정확도를 기록하며 모델의 실제 세계 인식 능력 한계를 드러낸다.

핵심 기여도

핵심 아이디어

기존 MLLM 평가 벤치마크는 데이터 규모가 작고, 라벨 품질이 낮으며, 실제 세계 시나리오의 복잡성과 고해상도 이미지를 반영하지 못한다는 문제점을 지적한다. 이를 해결하기 위해, MME-RealWorld는 고해상도 이미지를 기반으로 실제 세계 시나리오(예: 자율주행, 감시 영상, 원격 감지)를 반영한 43개의 하위 태스크를 구성했다. 라벨링 과정에서는 25명의 전문 라벨러와 7명의 MLLM 전문가가 참여해 수작업으로 QA 쌍을 생성함으로써 높은 품질을 보장했다. 특히, 대부분의 질문이 인간에게도 어려운 수준으로 설계되어, MLLM의 진정한 인식 능력을 평가할 수 있도록 했다.

기술적 접근법

주요 결과

의의 및 한계

MME-RealWorld는 기존 벤치마크의 데이터 규모, 라벨 품질, 태스크 난이도 문제를 해결한 최대 규모의 수작업 라벨 고해상도 벤치마크로, MLLM의 실제 세계 인식 능력을 정확히 평가할 수 있는 기반이 된다. 특히, 중국어 시나리오를 반영한 MME-RealWorld-CN은 언어 및 문화적 맥락에 따른 모델 성능 차이를 분석하는 데 유용하다. 그러나, 라벨링 과정에서의 비용과 시간이 매우 높아 대규모 확장에는 한계가 있을 수 있다.

실용적 활용

MME-RealWorld는 자율주행, 감시 시스템, 원격 감지 등 고해상도 이미지 처리가 필요한 산업 분야에서 MLLM의 성능 평가에 활용될 수 있다. 또한, 중국어 사용 환경에서의 모델 성능 비교 및 최적화에도 유용하며, MLLM 연구자들이 실제 세계 시나리오에 대한 모델 개선 방향을 도출하는 데 기여할 수 있다.