Agent-as-a-Judge: Evaluate Agents with Agents

Mingchen Zhuge, Changsheng Zhao, Dylan R. Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, Jurgen Schmidhuber

arXiv:2410.10934 · 2026-07-27 공개 · arXiv · PDF

code-generation llm-as-a-judge evaluation-framework agentic-systems self-improvement agent-as-a-judge devai-benchmark reward-signals

Abstract

Contemporary evaluation techniques are inadequate for agentic systems. These approaches either focus exclusively on final outcomes -- ignoring the step-by-step nature of agentic systems, or require excessive manual labour. To address this, we introduce the Agent-as-a-Judge framework, wherein agentic systems are used to evaluate agentic systems. This is an organic extension of the LLM-as-a-Judge framework, incorporating agentic features that enable intermediate feedback for the entire task-solving process. We apply the Agent-as-a-Judge to the task of code generation. To overcome issues with existing benchmarks and provide a proof-of-concept testbed for Agent-as-a-Judge, we present DevAI, a new benchmark of 55 realistic automated AI development tasks. It includes rich manual annotations, like a total of 365 hierarchical user requirements. We benchmark three of the popular agentic systems using Agent-as-a-Judge and find it dramatically outperforms LLM-as-a-Judge and is as reliable as our human evaluation baseline. Altogether, we believe that Agent-as-a-Judge marks a concrete step forward for modern agentic systems -- by providing rich and reliable reward signals necessary for dynamic and scalable self-improvement.

한국어 요약

한 줄 요약

Agent-as-a-Judge를 제안하며, DevAI 벤치마크에서 90%의 인간 평가자와 일치율을 보임.

핵심 기여도

핵심 아이디어

기존 평가 방법은 agentic 시스템의 단계적 문제 해결 과정을 무시하거나, 과도한 수작업을 요구한다. 이에 반해, Agent-as-a-Judge는 agentic 시스템을 사용하여 agentic 시스템을 평가하는 새로운 프레임워크로, LLM-as-a-Judge의 확장형이다. 이는 agentic 시스템이 인간처럼 단계적으로 사고하고 행동하는 특성을 반영하여, 전체 과정에서 중간 피드백을 제공할 수 있도록 설계되었다. DevAI 벤치마크는 기존 코드 생성 평가셋(예: HumanEval, MBPP)이 실용적 문제를 반영하지 못한다는 점을 개선하기 위해 55개의 실제 AI 개발 태스크와 365개의 계층적 요구사항을 포함한 데이터셋으로 구성되었다.

기술적 접근법

주요 결과

의의 및 한계

Agent-as-a-Judge는 agentic 시스템의 단계적 평가를 가능하게 하며, 인간 평가자와 유사한 신뢰도를 제공함으로써, agentic 시스템의 자기 개선과 확장에 기여할 수 있다. DevAI는 기존 코드 생성 평가셋이 실용성을 반영하지 못하는 문제를 해결하고, agentic 시스템의 종합적 평가를 위한 기반을 제공한다. 그러나, Agent-as-a-Judge는 아직 개념 증명 단계에 머무르며, 다양한 도메인에서의 일반화 가능성과 장기적 신뢰도 검증이 필요하다. 또한, 인간 평가자와의 의견 차이가 발생할 경우, 추가적인 토론 라운드나 전문가 패널이 필요하다는 한계가 있다.

실용적 활용

Agent-as-a-Judge는 소프트웨어 개발, 자동화 테스트, AI 기반 제품 개발 등에서 agentic 시스템의 효율적 평가를 지원할 수 있다. DevAI는 AI 개발자 도구의 성능 검증 및 비교에 활용 가능하며, 기업이 agentic 시스템의 실제 작업 능력을 평가하는 데 유용한 벤치마크로 사용될 수 있다.