MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark

Dongping Chen, Ruoxi Chen, Shilin Zhang, Yinuo Liu, Yaochen Wang, Huichi Zhou, Qihui Zhang, Pan Zhou, Yao Wan, Lichao Sun

arXiv:2402.04788 · 2026-07-27 공개 · arXiv · PDF

vision-language multimodal-llm llm-bias hallucination mllm-as-a-judge scoring-evaluation pair-comparison batch-ranking

Abstract

Multimodal Large Language Models (MLLMs) have gained significant attention recently, showing remarkable potential in artificial general intelligence. However, assessing the utility of MLLMs presents considerable challenges, primarily due to the absence of multimodal benchmarks that align with human preferences. Drawing inspiration from the concept of LLM-as-a-Judge within LLMs, this paper introduces a novel benchmark, termed MLLM-as-a-Judge, to assess the ability of MLLMs in assisting judges across diverse modalities, encompassing three distinct tasks: Scoring Evaluation, Pair Comparison, and Batch Ranking. Our study reveals that, while MLLMs demonstrate remarkable human-like discernment in Pair Comparison, there is a significant divergence from human preferences in Scoring Evaluation and Batch Ranking. Furthermore, a closer examination reveals persistent challenges in the judgment capacities of LLMs, including diverse biases, hallucinatory responses, and inconsistencies in judgment, even in advanced models such as GPT-4V. These findings emphasize the pressing need for enhancements and further research efforts to be undertaken before regarding MLLMs as fully reliable evaluators. In light of this, we advocate for additional efforts dedicated to supporting the continuous development within the domain of MLLM functioning as judges. The code and dataset are publicly available at our project homepage: \url{https://mllm-judge.github.io/}.

한국어 요약

한 줄 요약

MLLM-as-a-Judge 벤치마크를 통해 MLLM의 판단 능력을 평가하고, 인간 기준과의 차이를 분석한다.

핵심 기여도

핵심 아이디어

기존 LLM-as-a-Judge 개념을 다중 모달로 확장한 MLLM-as-a-Judge 벤치마크를 제안함. 이는 MLLM이 인간과 얼마나 유사한 판단을 할 수 있는지를 평가하기 위한 체계적인 프레임워크를 제공한다. 특히, Pair Comparison에서는 GPT-4V가 높은 인간 유사도를 보였지만, Scoring Evaluation과 Batch Ranking에서는 MLLM들이 인간 기준과 큰 차이를 보임. 이는 MLLM이 단순한 비교보다는 정량적 평가나 다중 항목 순위 매기기에서 어려움을 겪는다는 점을 시사한다. 연구는 CoT와 Vision Expert System를 도입함으로써 판단 편향과 hallucination을 줄이는 방안을 제시함.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 MLLM을 인간과 유사한 판단을 할 수 있는 평가자로 사용할 수 있는지에 대한 체계적인 평가를 제공하며, MLLM의 판단 능력 한계를 명확히 드러냄. 특히, Pair Comparison에서의 성능은 MLLM의 잠재력을 보여주지만, Scoring Evaluation과 Batch Ranking에서의 부족함은 MLLM을 신뢰할 수 있는 평가자로 삼기에는 아직 한계가 있음을 시사함. 한계로는 MLLM의 내재적 편향과 hallucination 문제가 지적되며, CoT나 Vision Expert System을 도입하는 등의 기법이 필요함.

실용적 활용

MLLM-as-a-Judge는 디지털 콘텐츠 생성, 자동 평가 시스템, AI 기반 검토 도구 등에서 활용될 수 있음. 특히, 인간 평가자 대체나 대규모 평가 자동화에 유용할 수 있으며, CoT와 Vision Expert System을 결합한 방식은 MLLM의 신뢰도를 높이는 데 기여할 수 있음.