AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

Polina Kirichenko, Mark Ibrahim, Kamalika Chaudhuri, Samuel J. Bell

arXiv:2506.09038 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation model-scaling llm-reliability system-prompt benchmarking-llms uncertainty-reasoning abstention-bench abstention-degradation

Abstract

For Large Language Models (LLMs) to be reliably deployed in both everyday and high-stakes domains, knowing when not to answer is equally critical as answering correctly. Real-world user queries, which can be underspecified, ill-posed, or fundamentally unanswerable, require LLMs to reason about uncertainty and selectively abstain -- i.e., refuse to answer definitively. However, abstention remains understudied, without a systematic evaluation framework for modern LLMs. In this work, we introduce AbstentionBench, a large-scale benchmark for holistically evaluating abstention across 20 diverse datasets, including questions with unknown answers, underspecification, false premises, subjective interpretations, and outdated information. Evaluating 20 frontier LLMs reveals abstention is an unsolved problem, and one where scaling models is of little use. While recent reasoning LLMs have shown impressive results in complex problem solving, surprisingly, we find that reasoning fine-tuning degrades abstention (by $24\%$ on average), even for math and science domains on which reasoning models are explicitly trained. We find that while a carefully crafted system prompt can boost abstention in practice, it does not resolve models' fundamental inability to reason about uncertainty. We release AbstentionBench to foster research into advancing LLM reliability.

한국어 요약

한 줄 요약

AbstentionBench는 LLM이 불확실한 질문에 대해 답변을 거절하는 능력을 평가하는 대규모 벤치마크로, 추론 훈련이 거절 능력을 오히려 악화시킨다는 점을 밝혀낸다.

핵심 기여도

핵심 아이디어

LLM이 신뢰성 있게 사용되기 위해서는 "언제 답변하지 않을지" 판단하는 능력이 필수적이다. 그러나 기존 연구는 주로 정확성에 집중했으며, 거절 능력은 체계적으로 평가되지 않았다. AbstentionBench는 이 문제를 해결하기 위해 설계된 벤치마크로, 6가지 주요 시나리오(예: 정보 부족, 주관적 해석, 오래된 정보 등)를 포함한다. 특히, 추론 훈련이 정확성은 높이지만, 거절 능력을 오히려 악화시킨다는 점은 기존 인식과 반대되는 통찰이다. 이는 추론 모델이 정답을 강하게 추구하는 경향이 거절 판단을 방해한다는 의미이다.

기술적 접근법

주요 결과

의의 및 한계

AbstentionBench는 LLM의 거절 능력 평가에 체계적인 기반을 제공하며, 추론 훈련이 신뢰성 향상에 부정적 영향을 미칠 수 있음을 밝혀내는 학술적 의의가 있다. 그러나 현재 시스템 프롬프트는 근본적 문제 해결에 한계가 있으며, 새로운 훈련 방법이 필요하다는 점이 한계로 지적된다. 또한, 추론 훈련이 정확성은 높이지만, 불확실성 판단 능력을 약화시킨다는 점은 추후 연구 주제로 제시된다.

실용적 활용

AbstentionBench는 의료, 법률, 금융 등 고위험 분야에서 LLM의 신뢰성을 향상시키는 데 활용 가능하다. 특히, 모델이 정답을 강하게 추구하는 경향을 조정하여, 사용자에게 불확실성을 명확히 전달하는 시스템 설계에 기여할 수 있다.