Evaluation and Benchmarking of LLM Agents: A Survey

Mahmoud Mohammadi, Yipeng Li, Jean-Pierre Lo, W. Yip

arXiv:2507.21504 · 2026-07-27 공개 · arXiv · PDF

llm-agents benchmarking agent-reliability real-world-deployment dynamic-interactions evaluation-taxonomy enterprise-challenges metric-computation

Abstract

The rise of LLM-based agents has opened new frontiers in AI applications, yet evaluating these agents remains a complex and underdeveloped area. This survey provides an in-depth overview of the emerging field of LLM agent evaluation, introducing a two-dimensional taxonomy that organizes existing work along (1) evaluation objectives-what to evaluate, such as agent behavior, capabilities, reliability, and safety-and (2) evaluation process-how to evaluate, including interaction modes, datasets and benchmarks, metric computation methods, and tooling. In addition to taxonomy, we highlight enterprise-specific challenges, such as role-based access to data, the need for reliability guarantees, dynamic and long-horizon interactions, and compliance, which are often overlooked in current research. We also identify the future research directions, including holistic, more realistic, and scalable evaluation. This work aims to bring clarity to the fragmented landscape of agent evaluation and provide a framework for systematic assessment, enabling researchers and practitioners to evaluate LLM agents for real-world deployment.

한국어 요약

한 줄 요약

LLM 에이전트 평가의 현황과 주요 과제를 체계적으로 정리한 서베이 논문.

핵심 기여도

핵심 아이디어

LLM 에이전트는 단순한 언어 생성을 넘어, 도구 사용, 기억, 협업, 장기 계획 등 복합적 행동을 수행하므로, 기존 LLM 평가 방법론과는 차별화된 접근이 필요하다.
기존 연구는 특정 능력(예: 추론, 도구 사용)에만 초점을 맞추거나, 평가 목적과 과정을 통합적으로 다루지 못했다.
본 논문은 **평가 목적**(what to evaluate)과 **평가 과정**(how to evaluate)이라는 2차원 페러다임을 도입하여, 에이전트 평가의 복잡성을 체계적으로 분류하고, 기업 적용 시의 특수성을 반영한 평가 전략을 제시한다.
이러한 접근은 에이전트가 실제 환경에서 어떻게 작동하는지를 이해하고, 신뢰성과 안전성을 확보하는 데 기여한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 LLM 에이전트 평가의 분산된 연구들을 체계적으로 정리하고, 평가 목적과 과정을 구조화함으로써 연구자와 실무자에게 명확한 지침을 제공한다.
특히 기업 환경에서의 평가 요구사항(역할 기반 접근, 컴플라이언스)을 강조하며, 기존 연구에서 간과된 실용적 측면을 보완한다.
그러나, 평가 지표의 표준화가 부족하며, 다양한 도구와 벤치마크 간의 비교 기준이 명확하지 않은 점은 한계로 지적된다.
또한, 실제 배포 환경에서의 평가가 여전히 어려운 점도 언급된다.

실용적 활용

LLM 에이전트를 고객 서비스, 코드 생성, 디지털 어시스턴트 등에 적용하는 기업은 본 연구의 평가 틀을 활용하여 에이전트의 신뢰성과 안전성을 체계적으로 검증할 수 있다.
또한, 평가 도구(예: DeepEval, AgentOps)를 개발 환경에 통합하여 지속적인 모니터링과 개선을 지원할 수 있다.
연구자들은 제시된 2차원 분류법을 기반으로 새로운 평가 방법론을 설계할 수 있으며, 기업 맞춤형 평가 프레임워크를 개발할 수 있다.