AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation

Giulio Zeloni, Enrico Lo Conte, Salvatore Rionero, Giuseppe Santoro, Alessandro Rastelli, Fabio Sorrentino

arXiv:2610.01218 · 2026-10-07 공개 · arXiv · PDF

retrieval-augmented-generation llm-judge meta-evaluation beta-binomial decision-model stratified-evaluation risk-quantification quality-gate

Abstract

Enterprises adopting retrieval-augmented generation (RAG) face a recurring operational decision: promote, revise, or block a system version. The evidence is incomplete and the metrics come from fallible LLM judges. We report on AGO AI Quality Gate (AGO), an evidence-first quality-gate framework deployed in industrial RAG assessment engagements. AGO integrates four key components: a four-state decision model that treats missing data and judge errors as explicit outcomes; layered scoring combining deterministic checks, local guardrails, and structured LLM evaluation; a stratified beta-binomial gate that quantifies regression risk probabilistically; and a mandatory meta-evaluation protocol to validate the LLM judge before it influences decisions. Since engagement data is proprietary, we evaluate the judge layer on RAGBench, a public benchmark of 100k annotated RAG traces across 12 datasets. On identical stratified test samples (N=1200 per judge), a low-cost judge (gpt-4.1-nano) detects non-adherent answers barely above chance (AUROC 0.603 [0.570, 0.634]), despite producing flawless protocol output, while gpt-4o reaches 0.783 [0.756, 0.807] -- yet its per-domain performance still ranges from 0.62 to 0.88. A fixed-seed gate study spanning regression, no change, and improvement quantifies unsafe promotion, false-alarm cost, and improvement throughput. Under regression, the decision-grade profile reduces unsafe promotion to 22.2%-35.1%, against 29.3%-41.8% for a naive gate. These results support the design choices that judge quality must be measured per engagement and that point estimates alone are not a release decision.

한국어 요약

한 줄 요약

AGO AI Quality Gate는 RAG 시스템의 출시 결정을 위한 증거 중심 품질 게이트 프레임워크로, 판단자(LLM)의 신뢰도와 통계적 불확실성을 명시적으로 고려한다.

핵심 기여도

핵심 아이디어

AGO는 기업이 RAG 시스템의 버전을 출시할지, 수정할지, 차단할지를 결정하는 데 사용되는 품질 게이트 프레임워크이다. 기존 평가 방식은 불완전한 증거와 신뢰도가 낮은 LLM 판단자에 기반하여 오류를 내포하고 있다. AGO는 이 문제를 해결하기 위해 판단자의 신뢰도를 사전에 평가하고, 통계적 불확실성을 명시적으로 반영하는 4단계 결정 모델을 도입한다. 특히, 베타-이항 게이트는 각 범주(stratum)별 신뢰 구간을 계산하여 회귀 리스크를 확률적으로 표현하며, 이는 기존의 점 추정(point estimate) 방식보다 더 신뢰할 수 있는 결정을 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

AGO는 RAG 평가를 점수 생성에서 감사 가능한 출시 결정으로 전환시킨다. 판단자의 신뢰도를 사전에 평가하고, 통계적 불확실성을 명시적으로 반영함으로써 기존의 점 추정 방식의 한계를 극복한다. 특히, 판단자가 완벽한 프로토콜 출력을 제공하더라도 실제 구분 능력이 낮을 수 있음을 보여주는 실험 결과는 판단자 검증의 중요성을 강조한다. 한편, 기업 데이터는 비공개이므로 완전한 엔드-투-엔드 평가가 제한되며, 판단자 신뢰도는 도메인별로 달라져 일반화가 어렵다.

실용적 활용

AGO는 금융, 의료, 컴플라이언스 등 민감한 문서 기반 시스템에서 RAG 버전 출시 결정을 지원할 수 있다. 특히, 판단자의 신뢰도를 사전에 평가하고, 통계적 불확실성을 명시적으로 반영하는 방식은 규제 준수와 책임성 있는 AI 운영에 유용하다.