llm-evaluation model-scaling llm-reliability system-prompt benchmarking-llms uncertainty-reasoning abstention-bench abstention-degradation
Abstract
For Large Language Models (LLMs) to be reliably deployed in both everyday and high-stakes domains, knowing when not to answer is equally critical as answering correctly. Real-world user queries, which can be underspecified, ill-posed, or fundamentally unanswerable, require LLMs to reason about uncertainty and selectively abstain -- i.e., refuse to answer definitively. However, abstention remains understudied, without a systematic evaluation framework for modern LLMs. In this work, we introduce AbstentionBench, a large-scale benchmark for holistically evaluating abstention across 20 diverse datasets, including questions with unknown answers, underspecification, false premises, subjective interpretations, and outdated information. Evaluating 20 frontier LLMs reveals abstention is an unsolved problem, and one where scaling models is of little use. While recent reasoning LLMs have shown impressive results in complex problem solving, surprisingly, we find that reasoning fine-tuning degrades abstention (by $24\%$ on average), even for math and science domains on which reasoning models are explicitly trained. We find that while a carefully crafted system prompt can boost abstention in practice, it does not resolve models' fundamental inability to reason about uncertainty. We release AbstentionBench to foster research into advancing LLM reliability.
한국어 요약
한 줄 요약
AbstentionBench는 LLM이 불확실한 질문에 대해 답변을 거절하는 능력을 평가하는 대규모 벤치마크로, 추론 훈련이 거절 능력을 오히려 악화시킨다는 점을 밝혀낸다.
핵심 기여도
- AbstentionBench: 20개 다각적 데이터셋을 활용한 거절 능력 평가 벤치마크 도입.
- 20개 최첨단 LLM 평가 결과, 모델 확장이 거절 능력 개선에 효과 없음.
- 추론 훈련이 거절 능력을 평균 24% 저하시키는 것으로 밝혀짐.
- 시스템 프롬프트로 일시적 개선 가능하나, 근본적 문제 해결 불가.
핵심 아이디어
LLM이 신뢰성 있게 사용되기 위해서는 "언제 답변하지 않을지" 판단하는 능력이 필수적이다. 그러나 기존 연구는 주로 정확성에 집중했으며, 거절 능력은 체계적으로 평가되지 않았다. AbstentionBench는 이 문제를 해결하기 위해 설계된 벤치마크로, 6가지 주요 시나리오(예: 정보 부족, 주관적 해석, 오래된 정보 등)를 포함한다. 특히, 추론 훈련이 정확성은 높이지만, 거절 능력을 오히려 악화시킨다는 점은 기존 인식과 반대되는 통찰이다. 이는 추론 모델이 정답을 강하게 추구하는 경향이 거절 판단을 방해한다는 의미이다.
기술적 접근법
- **AbstentionBench**: 17개 기존 데이터셋 + 3개 추론 중심 벤치마크(GSM8K-Abstain, GPQA-Abstain, MMLU-Abstain)로 구성.
- **평가 모델**: 20개 최첨단 LLM (개방형 및 폐쇄형 포함).
- **자동 평가**: 품질 검증된 LLM 판별기 사용.
- **추론 훈련 실험**: DeepSeek R1 (Distill Llama 70B), s1 등 추론 모델과 비추론 모델 비교.
- **시스템 프롬프트 실험**: 특정 프롬프트로 거절률을 일시적으로 향상시키는 실험 수행.
주요 결과
- 20개 최첨단 LLM 평가 결과, 거절 능력은 대부분 저하됨.
- 추론 훈련 모델(예: DeepSeek R1)은 비추론 모델 대비 거절률 평균 24% 감소.
- 추론 토큰 예산 확장은 정확성은 향상시키나, 거절 능력은 악화.
- 시스템 프롬프트는 거절률을 일정 수준 개선하지만, 근본적 문제 해결은 어려움.
의의 및 한계
AbstentionBench는 LLM의 거절 능력 평가에 체계적인 기반을 제공하며, 추론 훈련이 신뢰성 향상에 부정적 영향을 미칠 수 있음을 밝혀내는 학술적 의의가 있다. 그러나 현재 시스템 프롬프트는 근본적 문제 해결에 한계가 있으며, 새로운 훈련 방법이 필요하다는 점이 한계로 지적된다. 또한, 추론 훈련이 정확성은 높이지만, 불확실성 판단 능력을 약화시킨다는 점은 추후 연구 주제로 제시된다.
실용적 활용
AbstentionBench는 의료, 법률, 금융 등 고위험 분야에서 LLM의 신뢰성을 향상시키는 데 활용 가능하다. 특히, 모델이 정답을 강하게 추구하는 경향을 조정하여, 사용자에게 불확실성을 명확히 전달하는 시스템 설계에 기여할 수 있다.