個人化幻象:大型語言模型如何捏造用戶畫像,以及為何自我監控具有誤導性
The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads
August 5, 2026
作者: Yushi Sun, Yanjie Zhang, Rui Sheng
cs.AI
摘要
具有持久記憶的個人化大型語言模型日益普及,然而其使用者模型的忠實度仍未受到檢驗。我們研究過度推論(OI):即大型語言模型在證據支持範圍之外捏造使用者屬性的現象。我們提出 MirageBench,包含:均衡分布於刻板、反刻板與中性輪廓的 150 個人物原型;涵蓋「想像梯度」的 6 項個人化任務;由獨立評判者操作化的四分類忠實度分類法(該評判者在 400 條斷言上對照一位盲測人工標註者進行驗證,四分類的柯恩卡帕係數 = 0.863,二元分類的卡帕係數 = 0.900);以及由 12 個模型(分屬 7 個家族)在 143,616 條經評判斷言上組成的排行榜。我們發現過度推論普遍存在:在此評估中,12 個模型無一例外,每個模型都在 35% 至 49% 的斷言上出現過度推論(跨模型平均值為 41.6%;以斷言加權後為 41.8%)。最引人注目的是,我們揭示了一種「自我監控反轉」:在模型選擇層面上,模型自我評估的過度推論與其經評判者測量的過度推論呈負等級相關(rho = -0.60,p = 0.044;探索性結果,自助法信賴區間寬廣,介於 [-0.90, +0.06],n = 12)。那些自我報告過度推論最少的模型,往往被標記為捏造最多的模型;因此,在比較模型時,自我報告的信心是誤導性的訊號,即使在單一模型內部,自我審查仍能對該模型自身的斷言進行中等程度的排序(AUROC 介於 0.58 至 0.83)。我們進一步顯示,過度推論具有任務依賴性(27% 至 59%),而且在多輪試驗中,推論出的屬性會近似線性地累積,且幾乎沒有修正。MirageBench 將外部驗證(而非模型自我報告)定位為實現值得信賴的個人化時更可靠的基礎。
English
Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.