개인화의 신기루: LLM이 사용자 프로필을 어떻게 조작하는가, 그리고 자기 모니터링이 오도하는 이유
The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads
August 5, 2026
저자: Yushi Sun, Yanjie Zhang, Rui Sheng
cs.AI
초록
지속적 메모리를 갖춘 개인화된 LLM(대규모 언어 모델)이 점점 더 배포되고 있지만, 그 사용자 모델의 충실성은 여전히 검토되지 않은 상태이다. 우리는 과잉 추론(over-inference, OI), 즉 LLM이 증거가 뒷받침하는 수준을 넘어 사용자 속성을 날조하는 현상을 연구한다. 우리는 MirageBench를 소개한다. 이는 고정관념적, 반고정관념적, 중립적 프로필에 걸쳐 균형 있게 구성된 150개의 페르소나, '상상력 연속체'를 포괄하는 6가지 개인화 작업, 독립 판정자에 의해 조작화된 4분류 충실성 분류체계(400개 주장에 대한 맹검 인간 주석자와의 대조 검증에서 코헨의 카파 = 0.863(4분류), 카파 = 0.900(이분 분류)), 그리고 판정된 143,616개 주장에 대한 7개 계열에 걸친 12개 모델의 리더보드로 구성된다. 우리는 과잉 추론이 만연함을 발견한다. 평가에 포함된 12개 모델 각각은 자기 주장의 35%~49%에서 과잉 추론을 보였고(모델 간 평균 41.6%, 주장 가중 평균 41.8%), 어떤 모델도 예외가 아니었다. 가장 두드러진 발견은 자기 모니터링 역전(Self-Monitoring Inversion)이다. 모델 선택 수준에서 모델들이 스스로 평가한 OI는 판정자 측정 OI와 음의 순위 상관을 보인다(rho = -0.60, p = 0.044; 탐색적 분석, 넓은 부트스트랩 신뢰구간 [-0.90, +0.06], n = 12). 과잉 추론을 가장 적게 보고하는 모델일수록 가장 많이 날조하는 것으로 지적되는 경향이 있으므로, 단일 모델 내에서 자체 감사가 해당 모델 자신의 주장들을 여전히 중간 수준으로 잘 순위화한다 하더라도(AUROC 0.58~0.83), 자기 보고된 확신도는 모델 비교에 있어 오해를 부르는 신호이다. 우리는 또한 OI가 작업 의존적이며(27%~59%), 다중 턴 예비 실험에서는 추론된 속성이 거의 수정 없이 거의 선형적으로 축적됨을 보인다. MirageBench는 모델 자기 보고보다 외부 검증을 신뢰할 수 있는 개인화를 위한 더 신뢰할 만한 기반으로 제시한다.
English
Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.