LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

A. Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, R. Fernández, Albert Gatt, E. Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, Andr'e F. T. Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K Surikuchi, Ece Takmaz, A. Testoni

arXiv:2406.18403 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation benchmark-datasets llm-reliability human-judgment evaluation-validity nlp-annotation judgment-replication model-generated-text

Abstract

There is an increasing trend towards evaluating NLP models with LLMs instead of human judgments, raising questions about the validity of these evaluations, as well as their reproducibility in the case of proprietary models. We provide JUDGE-BENCH, an extensible collection of 20 NLP datasets with human annotations covering a broad range of evaluated properties and types of data, and comprehensively evaluate 11 current LLMs, covering both open-weight and proprietary models, for their ability to replicate the annotations. Our evaluations show substantial variance across models and datasets. Models are reliable evaluators on some tasks, but overall display substantial variability depending on the property being evaluated, the expertise level of the human judges, and whether the language is human or model-generated. We conclude that LLMs should be carefully validated against human judgments before being used as evaluators.

한국어 요약

한 줄 요약

LLM을 인간 평가자 대체로 사용할 수 있는지 20개 NLP 평가 태스크에서 대규모 실험을 통해 검증한 연구.

핵심 기여도

핵심 아이디어

LLM을 NLP 모델 평가 도구로 사용하는 경향이 증가하고 있지만, 이는 인간 평가와의 일치도, 재현 가능성, 편향 가능성 등 여러 측면에서 의문을 제기한다. 본 연구는 20개의 다양한 NLP 평가 태스크와 평가 속성(예: 일관성, 유창성, 독성 등)을 포함한 Judge-Bench를 통해 LLM 평가의 신뢰도를 대규모로 평가했다. 핵심 통찰은 LLM이 특정 태스크(예: 지시 준수)에서는 신뢰할 수 있지만, 평가 속성, 평가자 전문성, 데이터 유형에 따라 결과가 크게 변한다는 점이다. 특히, 독성 평가와 같은 복잡한 속성에서는 LLM이 인간 평가와의 일치도가 낮고, 응답률도 저조한 것으로 나타났다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 LLM을 평가 도구로 사용할 때의 한계를 명확히 보여주며, 평가 전략의 신중한 검증이 필요함을 강조한다. 특히, 평가 속성과 데이터 유형에 따라 LLM의 신뢰도가 달라지므로, 단일 모델을 다양한 평가에 적용하는 것은 위험할 수 있다. 한계점으로는 일부 데이터셋(예: QAGS, Recipe-generation, NewsRoom)에서는 모델 점수가 인간 상한선에 크게 못 미치는 점이 지적된다. 또한, 평가 전략의 효과(예: Chain-of-Thought)가 일관되지 않아 추가 연구가 필요하다.

실용적 활용

이 연구는 NLP 모델 평가 시 LLM을 도구로 활용하려는 연구자 및 엔지니어에게 중요한 참고 자료가 될 수 있다. 특히, 평가 전략을 선택할 때 평가 태스크와 속성에 따라 적절한 LLM을 선택하거나, 인간 평가와의 일치도를 사전에 검증하는 것이 중요하다. 이는 평가 결과의 신뢰도와 재현 가능성 향상에 기여할 수 있다.