評估大型語言模型中個性化的隱性成本
Evaluating the Hidden Costs of Personalization in Large Language Models
August 28, 2026
作者: Yumeng Wang, Yuchen Wu, Cheng Qian, Zhiyuan Fan, Hyeonjeong Ha, Shujin Wu, Jiayu Liu, Heng Ji, Ge Wang
cs.AI
摘要
雖然大型語言模型(LLMs)納入用戶個性化訊號以提升可用性與助益性,但當模型在個人化情境(如對話歷史、推斷的偏好及用戶檔案)的條件下進行運作時,其日益從提供平衡且具資訊性的回應,轉向最佳化用戶滿意度。具體而言,我們識別出三項新興風險:(1)無關個性化,即模型在不必要的語境中引用個人資訊;(2)偏好窄化,即模型強化資訊回聲室效應;以及(3)諂媚偏誤,即模型過度認同用戶觀點。因此,模型可能在不必要的語境中引用個人資訊,無意間導致回應多樣性下降,或過度迎合用戶意見。儘管個性化在AI助理中的應用日益普及,對其潛在副作用的系統性評估仍相當有限。為填補此研究缺口,我們提出PRISK,一個具動態評估框架,結合自動化資料生成與量身定制的指標,用以揭露當前LLM個性化的系統性限制,以及個性化資訊如何形塑其回應。我們對13個LLM的實證分析顯示,用戶檔案與檢索記憶的存在 consistently 加劇了各項偏誤,導致無關個性化平均下降45.9%、偏好窄化下降41.7%、諂媚偏誤下降61.7%。
English
While Large language models (LLMs) incorporate user personalization signals to improve usability and helpfulness, they increasingly shift from providing balanced, informative responses toward optimizing for user satisfaction when conditioned on personal context such as conversation history, inferred preferences, and user profiles. Specifically, we identify three emerging risks: (1) irrelevant personalization, where models reference personal information in unnecessary contexts; (2) preference narrowing, where models reinforce informational echo chambers; and (3) sycophantic bias, where models agree excessively with user opinions. As a result, models may reference personal information in contexts where it is unnecessary, inadvertently collapse response diversity, or agree excessively with user opinions. Despite the growing use of personalization in AI assistants, there has been limited systematic evaluation of its potential side effects. To bridge this gap, we propose PRISK, a dynamic evaluation framework with automated data generation and tailored metrics that uncovers systematic limitations in current LLM personalization and how personalized information shapes its responses. Our empirical analysis across 13 LLMs demonstrates the presence of user profiles and retrieved memories consistently exacerbates biases, resulting in an average drop of 45.9% in irrelevant personalization, 41.7% in preference narrowing and 61.7% in sycophantic bias.