Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict

Kaiser Sun, Bernal Jimenez Gutierrez, Hongjun Liu, Jingyu Zhang, Jie Gao, Mark Dredze, Daniel Khashabi

arXiv:2610.12360 · 2026-10-11 공개 · arXiv · PDF

llm-agents agent-harness trajectory-analysis model-interventions conflict-resolution task-accuracy knowledge-conflict epistemic-humility

Abstract

When retrieved evidence contradicts an agent's prior beliefs, does it revise its answer, acknowledge uncertainty, or persist with an incorrect conclusion? Existing evaluations of agentic systems focus primarily on task success, offering limited insight into how agents handle such conflicts. We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution. We operationalize EH through three trajectory-level behavioral dimensions: Identify, Solve, and Escalate (ISE). Through knowledge conflict, situations where the backbone language model's parametric knowledge contradicts the evidence it encounters, or where two contextual sources disagree, we evaluate two conflict settings: (1) controlled conflict and (2) naturally occurring conflict during multi-step agentic execution, each paired with matched no-conflict controls. Evaluating four agents, we find that higher task accuracy does not necessarily correspond to greater epistemic humility: some high-accuracy configurations recognize conflicts during execution but do not communicate unresolved uncertainty in their incorrect final answers. Trajectory-level analysis further reveals that agents frequently detect conflicts in early steps of execution but fail to maintain or resolve them in later steps. Finally, we show that model-level interventions can improve EH, but often at the cost of task accuracy, suggesting that epistemic humility emerges from the interaction among the backbone model, agent harness, and evaluation environment.

한국어 요약

한 줄 요약

LLM 에이전트가 지식 충돌 상황에서 겸손하게 불확실성을 인식하고 전달하는 능력을 평가하는 새로운 프레임워크를 제안한다.

핵심 기여도

핵심 아이디어

기존 평가가 주로 최종 정답 정확도에 집중하는 반면, 본 연구는 에이전트가 실행 과정에서 불확실성을 인식하고 전달하는 능력, 즉 **에피스테모틱 겸손(EH)**을 평가하는 새로운 프레임워크를 제안한다.
**ISE**는 에이전트가 (1) 충돌을 **Identify**하고, (2) 도구 사용으로 **Solve**하며, (3) 해결되지 않은 불확실성을 **Escalate**하는 행동을 측정한다.
이를 위해 **지식 충돌** 상황을 인위적으로 조성한 **Controlled conflict**와 실행 과정에서 자연스럽게 발생하는 **Naturally occurring conflict** 두 가지 설정을 사용한다.
실험 결과, **높은 정확도가 EH와 반드시 상관하지 않음**을 밝혀내며, **트래젝토리 수준의 평가**가 필요하다는 점을 강조한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 에이전트 평가에서 **최종 정답에만 집중하는 한계**를 지적하고, **실행 과정에서의 불확실성 처리 능력**을 평가하는 새로운 프레임워크를 제시한다.
**ISE**는 에이전트가 충돌을 인식하고, 해결하고, 전달하는 능력을 정량적으로 평가할 수 있는 구조를 제공하며, **에이전트 설계 및 평가 기준 개선**에 기여할 수 있다.
하지만, **모델, 하이버, 평가 환경 간 상호작용**의 영향을 분리하지 못한 점이 한계로 지적된다.
또한, **개입 효과**가 정확도에 부정적 영향을 미친다는 점도 주의 깊게 고려해야 할 문제점이다.

실용적 활용

본 연구는 **AI 에이전트가 사용자에게 신뢰성 있는 답변을 제공하기 위해 불확실성을 인식하고 전달하는 능력**을 평가하는 데 활용될 수 있다.
특히, **의료, 법률, 금융 등 고위험 분야**에서 에이전트가 **신뢰성 있는 정보를 전달**하거나 **불확실성을 명확히 전달**하는 능력을 평가하는 데 유용할 수 있다.
또한, **에이전트 훈련 과정에서 중간 단계 행동을 모니터링**하는 방식으로 **실질적인 EH 향상**을 유도할 수 있다.