个性化幻象:大语言模型如何编造用户画像,以及为何自我监控会误导
The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads
August 5, 2026
作者: Yushi Sun, Yanjie Zhang, Rui Sheng
cs.AI
摘要
具有持久记忆的个性化大语言模型(LLM)正被日益广泛地部署,但其用户模型的忠实性仍未得到检验。我们研究过度推断(OI):即大语言模型捏造超出证据支持范围的用户属性的现象。我们提出MirageBench,包含150个在刻板印象、反刻板印象与中性画像之间均衡分布的用户画像,6个跨越“想象梯度”的个性化任务,一个由独立评判器操作化的四分类忠实性分类体系(以盲法人类标注者为参照,在400条陈述上验证:四分类科恩卡帕系数=0.863,二分类卡帕=0.900),以及一个涵盖7个系列共12个模型、基于143,616条经评判陈述的排行榜。我们发现过度推断普遍存在:12个模型中的每一个都对其35%–49%的陈述进行过度推断(跨模型均值41.6%;按陈述加权41.8%),本评估中没有任何模型能够避免这一点。最引人注目的是,我们揭示了一种“自我监控反转”:在模型选择层面,模型的自我评估OI与其评判器测量OI呈负秩相关(rho = -0.60,p = 0.044;探索性结果,自助法置信区间较宽[-0.90, +0.06],n = 12)。报告过度推断最少的模型往往被标记为捏造最多的模型,因此自我报告的置信度在比较模型时是一种误导性信号,尽管在单个模型内部,自我审计仍能对该模型自身的陈述进行较好的排序(AUROC 0.58–0.83)。我们进一步表明,OI随任务不同而变化(27%–59%),并且在多轮试点中,推断出的属性近似线性累积,且很少得到修正。MirageBench将外部验证而非模型自我报告定位为可信个性化更可靠的基础。
English
Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.