Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

Irina Proskurina, Mayank Kumar, Oyindolapo O. Komolafe

arXiv:2608.13430 · 2026-08-16 공개 · arXiv · PDF

benchmark-evaluation language-models instruction-tuning question-answering model-calibration rationale-generation lexical-diversity confidence-measurement

Abstract

Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the consistency of the generated supporting rationales. In this paper, we study whether corresponding changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruction tuning consistently alters answer confidence, despite limited changes in predictive accuracy and decreases in likelihood-based calibration. Secondly, we observe a non-uniform effect of instruction tuning on rationale diversity: cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. Finally, we find that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning.

한국어 요약

한 줄 요약

지시 학습은 모델의 답변 신뢰도를 높이지만, 이는 일관된 어휘 다양성 변화와 관련되지 않음을 밝힘.

핵심 기여도

핵심 아이디어

지시 학습은 모델이 더 확신 있게 답변하도록 유도하지만, 이는 반드시 어휘 다양성의 일관된 변화와 연결되지 않는다. 이 연구는 `cross-rationale diversity`와 `surface-level lexical diversity`를 구분하여 분석함으로써, 지시 학습이 모델 신뢰도와 어휘 다양성에 미치는 영향이 복합적임을 밝혔다. `SelfBLEU` 지표를 사용하여 어휘 다양성을 측정하고, `H_choice(x)` 공식을 통해 선택된 답변에 대한 불확실성을 정량화함으로써, 신뢰도와 다양성 간의 관계를 정확히 분석했다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 지시 학습이 모델 신뢰도와 어휘 다양성에 미치는 영향이 복합적임을 밝혀, 신뢰도 추정 연구에 새로운 통찰을 제공한다. 특히, `cross-rationale diversity`와 `surface-level diversity`를 구분하여 분석함으로써, 단순한 어휘 다양성 측정이 모델 신뢰도를 정확히 반영하지 못함을 보여준다. 한계로는 `SelfBLEU`가 어휘적 다양성만 반영하고 의미적 다양성은 고려하지 못한다는 점이 언급된다.

실용적 활용

의료, 법률, 금융 분야에서 신뢰도 추정이 중요한 모델 개발에 활용 가능. 예를 들어, 모델이 답변에 대해 과도한 확신을 보일 때, 이 연구는 어휘 다양성 분석을 통해 신뢰도의 신뢰성을 평가하는 데 도움을 줄 수 있다.