Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, S. Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo

arXiv:2405.01535 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation language-models open-source human-alignment evaluation pairwise-ranking model-assessment direct-assessment

Abstract

Proprietary LMs such as GPT-4 are often employed to assess the quality of responses from various LMs. However, concerns including transparency, controllability, and affordability strongly motivate the development of open-source LMs specialized in evaluations. On the other hand, existing open evaluator LMs exhibit critical shortcomings: 1) they issue scores that significantly diverge from those assigned by humans, and 2) they lack the flexibility to perform both direct assessment and pairwise ranking, the two most prevalent forms of assessment. Additionally, they do not possess the ability to evaluate based on custom evaluation criteria, focusing instead on general attributes like helpfulness and harmlessness. To address these issues, we introduce Prometheus 2, a more powerful evaluator LM than its predecessor that closely mirrors human and GPT-4 judgements. Moreover, it is capable of processing both direct assessment and pair-wise ranking formats grouped with a user-defined evaluation criteria. On four direct assessment benchmarks and four pairwise ranking benchmarks, Prometheus 2 scores the highest correlation and agreement with humans and proprietary LM judges among all tested open evaluator LMs. Our models, code, and data are all publicly available.

한국어 요약

한 줄 요약

Prometheus 2는 직접 평가와 순위 평가 모두에서 GPT-4 수준의 성능을 보이는 오픈소스 평가 언어 모델이다.

핵심 기여도

핵심 아이디어

기존 오픈소스 평가 모델은 직접 평가와 순위 평가 중 하나만 수행하거나, 인간 평가와 GPT-4와의 일관성 부족 등의 문제를 보였다. Prometheus 2는 이 두 평가 방식을 모두 처리할 수 있도록 설계되었으며, 사용자 정의 평가 기준을 지원한다. 핵심 아이디어는 Feedback Collection(직접 평가)과 Preference Collection(순위 평가) 데이터셋에 기반한 두 개의 평가 모델을 학습한 후, 그 가중치를 병합하여 통합 모델을 생성하는 것이다. 이는 기존의 단일 형식 학습 모델보다 더 높은 유연성과 정확도를 제공한다.

기술적 접근법

주요 결과

의의 및 한계

Prometheus 2는 오픈소스 평가 모델로서 인간 및 GPT-4 평가와의 일관성을 높이며, 평가의 투명성과 제어 가능성을 향상시킨다. 특히, 사용자 정의 평가 기준을 지원함으로써 실제 산업 및 연구 환경에서의 유연한 적용이 가능하다. 그러나 모델의 학습 데이터셋이 특정 벤치마크에 의존적이라는 한계가 있으며, 더 다양한 도메인에서의 평가가 필요하다. 또한, 가중치 병합 방식의 내부 메커니즘에 대한 이론적 해석은 부족한 상태이다.

실용적 활용

Prometheus 2는 언어 모델의 질적 평가가 필요한 연구 및 산업 분야에서 활용 가능하다. 특히, 대규모 평가 과정에서 투명성과 비용 효율성을 요구하는 경우에 적합하며, 사용자 정의 기준을 기반으로 한 맞춤형 평가 시스템 구축에도 활용할 수 있다.