评估大语言模型中个性化的隐性成本
Evaluating the Hidden Costs of Personalization in Large Language Models
August 28, 2026
作者: Yumeng Wang, Yuchen Wu, Cheng Qian, Zhiyuan Fan, Hyeonjeong Ha, Shujin Wu, Jiayu Liu, Heng Ji, Ge Wang
cs.AI
摘要
尽管大语言模型(LLMs)通过整合用户个性化信号来提升可用性与帮助性,但当它们以对话历史、推断偏好和用户画像等个人情境为条件时,正日益从提供平衡、信息丰富的回复转向优化用户满意度。具体来说,我们识别出三种新兴风险:(1)无关个性化,即模型在不必要的情境下引用个人信息;(2)偏好窄化,即模型强化信息回音室;(3)谄媚偏差,即模型过度认同用户观点。因此,模型可能在不必要的情境中引用个人信息,无意中削弱响应多样性,或过度认同用户观点。尽管个性化在AI助手中的应用日益增长,但对其潜在副作用的系统性评估仍然有限。为弥补这一差距,我们提出了PRISK——一个结合自动化数据生成和定制化指标的动态评估框架,用于揭示当前LLM个性化中的系统性局限,以及个性化信息如何影响其响应。我们对13个LLM的实证分析表明,用户画像和检索到的记忆的存在持续加剧偏差,导致无关个性化平均下降45.9%,偏好窄化下降41.7%,谄媚偏差下降61.7%。
English
While Large language models (LLMs) incorporate user personalization signals to improve usability and helpfulness, they increasingly shift from providing balanced, informative responses toward optimizing for user satisfaction when conditioned on personal context such as conversation history, inferred preferences, and user profiles. Specifically, we identify three emerging risks: (1) irrelevant personalization, where models reference personal information in unnecessary contexts; (2) preference narrowing, where models reinforce informational echo chambers; and (3) sycophantic bias, where models agree excessively with user opinions. As a result, models may reference personal information in contexts where it is unnecessary, inadvertently collapse response diversity, or agree excessively with user opinions. Despite the growing use of personalization in AI assistants, there has been limited systematic evaluation of its potential side effects. To bridge this gap, we propose PRISK, a dynamic evaluation framework with automated data generation and tailored metrics that uncovers systematic limitations in current LLM personalization and how personalized information shapes its responses. Our empirical analysis across 13 LLMs demonstrates the presence of user profiles and retrieved memories consistently exacerbates biases, resulting in an average drop of 45.9% in irrelevant personalization, 41.7% in preference narrowing and 61.7% in sycophantic bias.