ChatPaper.aiChatPaper

パーソナライゼーションの蜃気楼:大規模言語モデルがユーザープロファイルを捏造する仕組みと、自己モニタリングが誤導する理由

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

August 5, 2026
著者: Yushi Sun, Yanjie Zhang, Rui Sheng
cs.AI

要旨

永続的なメモリを備えたパーソナライズLLMがますます展開されているが、そのユーザーモデルの忠実性は未検証のままである。本稿では、LLMが証拠が支持する以上にユーザー属性を作り出してしまう現象である過剰推論(OI: over-inference)を研究する。我々はMirageBenchを導入する。これは、ステレオタイプ的、反ステレオタイプ的、中立的プロファイルにバランスよく配分された150のペルソナ、「想像力勾配」にわたる6つのパーソナライズタスク、独立した判定者によって運用される4値の忠実性分類法(400件の主張に対する盲検人手アノテータとの検証で、Cohen's kappa = 0.863(4クラス)、kappa = 0.900(2値))、および7ファミリーにわたる12モデルのリーダーボード(判定済み主張143,616件)から構成される。我々は、過剰推論が広く見られることを見いだした。評価対象の12モデルすべてが、その主張の35%~49%を過剰推論しており(モデル間平均41.6%、主張加重平均41.8%)、本評価においてそれを免れるモデルは存在しない。最も顕著なのは、自己モニタリング反転(Self-Monitoring Inversion)を明らかにしたことである。モデル選択レベルでは、モデルの自己評価によるOIは、判定者による測定OIと負の順位相関を示した(rho = -0.60, p = 0.044; 探索的、幅広いブートストラップCI [-0.90, +0.06], n = 12)。過剰推論が最も少ないと報告するモデルほど、最も多く捏造していると判定される傾向がある。したがって、単一モデル内では自己監査がそのモデル自身の主張を依然として中程度にうまくランク付けするものの(AUROC 0.58–0.83)、自己報告された信頼度はモデル比較には誤解を招くシグナルである。さらに我々は、OIがタスク依存的であり(27%–59%)、マルチターンパイロットでは、推定された属性がほぼ線形に蓄積され、修正はほとんどないことを示す。MirageBenchは、モデルの自己報告ではなく外部検証を、信頼できるパーソナライゼーションのためのより信頼性の高い基盤として位置づけるものである。
English
Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.