Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael I. Jordan, Joseph E. Gonzalez, Ion Stoica

arXiv:2403.04132 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation pairwise-comparison human-preference open-platform crowdsourcing llm-leaderboard statistical-analysis

Abstract

Large Language Models (LLMs) have unlocked new capabilities and applications; however, evaluating the alignment with human preferences still poses significant challenges. To address this issue, we introduce Chatbot Arena, an open platform for evaluating LLMs based on human preferences. Our methodology employs a pairwise comparison approach and leverages input from a diverse user base through crowdsourcing. The platform has been operational for several months, amassing over 240K votes. This paper describes the platform, analyzes the data we have collected so far, and explains the tried-and-true statistical methods we are using for efficient and accurate evaluation and ranking of models. We confirm that the crowdsourced questions are sufficiently diverse and discriminating and that the crowdsourced human votes are in good agreement with those of expert raters. These analyses collectively establish a robust foundation for the credibility of Chatbot Arena. Because of its unique value and openness, Chatbot Arena has emerged as one of the most referenced LLM leaderboards, widely cited by leading LLM developers and companies. Our demo is publicly available at \url{https://chat.lmsys.org}.

한국어 요약

한 줄 요약

Chatbot Arena는 240K 이상의 사용자 투표를 바탕으로 LLM을 인간 선호도 기반으로 평가하는 오픈 플랫폼이다.

핵심 기여도

핵심 아이디어

Chatbot Arena는 기존 정적 벤치마크의 한계를 극복하기 위해 실시간 사용자 질문과 투표를 기반으로 LLM을 평가하는 새로운 접근법을 제시한다. 사용자는 두 익명의 LLM이 생성한 응답을 비교하고 선호도를 투표하며, 이 데이터는 E-values와 같은 통계 모델을 통해 신뢰성 있는 모델 랭킹을 도출한다. 이는 기존 정적 평가가 반영하지 못하는 유연성과 상호작용성을 측정할 수 있는 중요한 기회를 제공한다. 특히, 사용자 질문이 100개 이상의 언어로 수집되며, 전문가 평가와 높은 일관성을 보이는 점에서 데이터의 질이 입증된다.

기술적 접근법

주요 결과

의의 및 한계

Chatbot Arena는 LLM 평가에서 인간 선호도를 반영한 대규모 오픈 플랫폼으로, 기존 정적 평가 방식의 한계를 보완하는 데 기여한다. 특히, 사용자 질문의 다양성과 투표 데이터의 신뢰성을 입증하며, 산업 및 연구 분야에서 중요한 기준이 되고 있다. 그러나 사용자층이 주로 LLM 연구자와 아마추어 사용자로 구성되어 있어 편향 가능성은 존재한다. 또한, 안전성 평가를 포함하지 않아, 향후 안전성 측정 메커니즘의 개발이 필요하다.

실용적 활용

Chatbot Arena는 LLM 개발사, 연구소, 산업 현장에서 모델 성능을 실시간으로 비교하고 개선 방향을 도출하는 데 활용될 수 있다. 특히, 사용자 선호도 기반의 평가를 통해 실제 사용 환경에 더 가까운 모델 선택이 가능하다.