ChatPaper.aiChatPaper

정말 확신하시나요? 확실한가요? 지시 튜닝이 신뢰도와 어휘 다양성에 미치는 영향에 대하여

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

August 13, 2026
저자: Irina Proskurina, Mayank Kumar, Oyindolapo O. Komolafe
cs.AI

초록

지시 튜닝된 언어 모델은 다양한 생성 작업에서 우수한 성능을 달성하지만, 최근에는 언어화된 과신을 보이는 것으로도 나타났다. 질문 응답에서 모델의 언어화된 과신은 생성된 지지 근거의 일관성과 연관될 수 있다. 본 논문에서는 지시 튜닝으로 유도된 모델 신뢰도의 변화에 수반하여, 생성된 답변 근거의 어휘적 다양성에 상응하는 변화가 나타나는지 연구한다. 질문 응답 벤치마크들에 걸쳐 세 쌍의 매칭된 기본 및 지시 튜닝 모델을 평가한 결과, 예측 정확도의 제한된 변화와 우도 기반 캘리브레이션의 감소에도 불구하고, 지시 튜닝이 답변 신뢰도를 일관되게 변화시킨다는 것을 발견했다. 둘째, 지시 튜닝이 근거 다양성에 미치는 영향은 비균일함을 관찰했다: 근거 간 다양성은 일관되게 감소하는 반면, 표면적 어휘 다양성은 모델과 벤치마크에 따라 방향과 크기 모두에서 달라졌다. 마지막으로, 답변 선택과 근거 길이를 통제한 후에도 이러한 차이가 지속됨을 확인했으며, 이는 신뢰도와 근거 다양성이 지시 튜닝의 서로 다른 효과를 포착함을 보여준다.
English
Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the consistency of the generated supporting rationales. In this paper, we study whether corresponding changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruction tuning consistently alters answer confidence, despite limited changes in predictive accuracy and decreases in likelihood-based calibration. Secondly, we observe a non-uniform effect of instruction tuning on rationale diversity: cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. Finally, we find that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning.