ChatPaper.aiChatPaper

你确定你确定吗?论指令微调对置信度与词汇多样性的影响

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

August 13, 2026
作者: Irina Proskurina, Mayank Kumar, Oyindolapo O. Komolafe
cs.AI

摘要

指令微调的语言模型在一系列生成任务上表现出色,但近期研究也表明它们会展现出表述性的过度自信。在问答任务中,模型表述性的过度自信可能与其生成的支撑推理依据的一致性相关。本文研究了指令微调所引发的模型置信度变化是否伴随着生成的答案推理依据在词汇多样性上的相应变化。我们在问答基准上评估了三组匹配的基础模型和指令微调模型,发现指令微调持续地改变了答案置信度,尽管预测准确率变化有限,且基于似然的校准有所下降。其次,我们观察到指令微调对推理依据多样性的影响是非均匀的:跨推理依据的多样性持续下降,而表层词汇多样性在方向和幅度上均随模型和基准的不同而变化。最后,我们发现这些差异在控制答案选择和推理依据长度后仍然存在,从而证实置信度和推理依据多样性捕捉的是指令微调的不同效应。
English
Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the consistency of the generated supporting rationales. In this paper, we study whether corresponding changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruction tuning consistently alters answer confidence, despite limited changes in predictive accuracy and decreases in likelihood-based calibration. Secondly, we observe a non-uniform effect of instruction tuning on rationale diversity: cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. Finally, we find that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning.