ChatPaper.aiChatPaper

本当に確信していますか?指示チューニングが信頼度と語彙的多様性に与える影響について

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

August 13, 2026
著者: Irina Proskurina, Mayank Kumar, Oyindolapo O. Komolafe
cs.AI

要旨

指示チューニング済み言語モデルは、さまざまな生成タスクで高い性能を達成する一方で、近年、言語化された過信を示すことも明らかになっている。質問応答において、モデルの言語化された過信は、生成された根拠の一貫性と関連している可能性がある。本論文では、指示チューニングによって誘発されるモデルの信頼度の変化に伴って、生成される回答根拠の語彙的多様性に対応する変化が生じるかどうかを調査する。質問応答ベンチマークにおいてマッチした3組のベースモデルと指示チューニング済みモデルを評価し、予測精度の変化が限定的で、尤度に基づくキャリブレーションが低下するにもかかわらず、指示チューニングが回答の信頼度を一貫して変化させることを見いだした。次に、指示チューニングが根拠の多様性に及ぼす影響は一様ではないことを観察した。すなわち、根拠間の多様性は一貫して減少する一方、表層的な語彙の多様性はモデルとベンチマークに応じて方向と大きさの両方において異なる。最後に、これらの差異は回答選択と根拠の長さを制御した後も持続することを見いだし、信頼度と根拠の多様性が指示チューニングの異なる効果を捉えていることを確認する。
English
Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the consistency of the generated supporting rationales. In this paper, we study whether corresponding changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruction tuning consistently alters answer confidence, despite limited changes in predictive accuracy and decreases in likelihood-based calibration. Secondly, we observe a non-uniform effect of instruction tuning on rationale diversity: cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. Finally, we find that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning.