ChatPaper.aiChatPaper

你確定嗎?你真的確定嗎?論指令微調對信心與詞彙多樣性的影響

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

August 13, 2026
作者: Irina Proskurina, Mayank Kumar, Oyindolapo O. Komolafe
cs.AI

摘要

指令微調語言模型在各種生成任務上表現強勁,但近期也被發現會展現言語化的過度自信。在問答中,語言模型的言語化過度自信可能與其生成的支持性理據的一致性有關。本文研究了由指令微調引起的模型置信度變化,是否伴隨著生成答案理據的詞彙多樣性產生相應變化。我們在問答基準上評估了三組匹配的基礎模型與指令微調模型,發現指令微調一致地改變答案置信度,儘管預測準確率變化有限,且基於似然度的校準有所下降。其次,我們觀察到指令微調對理據多樣性的影響並非一致:跨理據多樣性持續下降,而表層詞彙多樣性在不同模型與基準之間的方向及幅度皆有差異。最後,我們發現這些差異在控制答案選擇與理據長度後仍然存在,證實置信度與理據多樣性捕捉到指令微調的不同效應。
English
Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the consistency of the generated supporting rationales. In this paper, we study whether corresponding changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruction tuning consistently alters answer confidence, despite limited changes in predictive accuracy and decreases in likelihood-based calibration. Secondly, we observe a non-uniform effect of instruction tuning on rationale diversity: cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. Finally, we find that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning.