Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey

Xiaoou Liu, Tiejin Chen, Longchao Da, Chacha Chen, Zhen-Yu Lin, Hua Wei

arXiv:2503.15850 · 2026-07-27 공개 · arXiv · PDF

large-language-models uncertainty-quantification confidence-calibration epistemic-uncertainty aleatoric-uncertainty scalable-uq input-ambiguity reasoning-divergence

Abstract

Uncertainty quantification (UQ) enhances the reliability of Large Language Models (LLMs) by estimating confidence in outputs, enabling risk mitigation and selective prediction. However, traditional UQ methods struggle with LLMs due to computational constraints and decoding inconsistencies. Moreover, LLMs introduce unique uncertainty sources, such as input ambiguity, reasoning path divergence, and decoding stochasticity, that extend beyond classical aleatoric and epistemic uncertainty. To address this, we introduce a new taxonomy that categorizes UQ methods based on computational efficiency and uncertainty dimensions, including input, reasoning, parameter, and prediction uncertainty. We evaluate existing techniques, summarize existing benchmarks and metrics for UQ, assess their real-world applicability, and identify open challenges, emphasizing the need for scalable, interpretable, and robust UQ approaches to enhance LLM reliability.

한국어 요약

한 줄 요약

LLM에서 불확실성 양화(UQ)의 중요성과 기존 방법의 한계를 분석하고, 새로운 분류체계를 제시한 조사 논문.

핵심 기여도

핵심 아이디어

LLM은 입력 모호성, 추론 경로 분기, 해독 불확실성 등 기존 ML 모델에서 다루지 못한 새로운 불확실성 요소를 포함한다. 이에 따라 기존 UQ 방법(예: Monte Carlo dropout, 앙상블)은 LLM에 적용 시 계산 비용이 높고 일관성이 부족하다. 본 논문은 LLM의 고유한 특성을 반영한 새로운 택서노미를 제시하며, UQ를 입력, 추론, 파라미터, 예측의 4가지 차원으로 구분한다. 각 차원은 aleatoric, epistemic, 또는 혼합 불확실성을 포함할 수 있다. 특히, 추론 불확실성은 기존 연구에서 다루어지지 않았으며, 이는 LLM의 복잡한 추론 과정에서 발생하는 문제를 반영한다.

기술적 접근법

주요 결과

의의 및 한계

본 논문은 LLM의 고유한 불확실성 구조를 반영한 체계적인 분류체계를 제시하며, UQ 연구의 새로운 방향성을 제시한다. 특히, 추론 불확실성과 같은 미개발 영역을 강조하며, 계산 효율성과 해석 가능성, 안정성 있는 UQ 접근법의 필요성을 강조한다. 그러나, LLM-as-a-judge 평가 방식은 편향과 불일치의 위험이 있으며, 인간 평가의 비용과 규모 제약도 한계로 지적된다.

실용적 활용

의료 진단, 법률 분석, 자율 운송 시스템 등 고위험 분야에서 LLM의 신뢰도를 향상시키는 데 활용 가능. 예를 들어, 의료 QA 시스템에서 불확실성 높은 진단은 전문가 검토로 전달되어 오진률을 감소시킬 수 있다.