ChatPaper.aiChatPaper

貌似合理但并非有效:对大语言模型作为合成调查受访者的心理测量学审计

Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents

July 6, 2026
作者: Mantas Lukauskas, Viktorija Šarkauskaitė
cs.AI

摘要

大型语言模型(LLMs)越来越多地被用作合成调查受访者,但现有评估只关注其回答在个体层面是否看似合理。我们认为,正确的问题应当是心理测量学层面的:LLMs 是否保留了真实人类调查数据的联合分布、潜在结构、信度、中介路径和人口学效应?我们引入了一个立陶宛组织心理学数据集(n=263 名员工;Dunham 变革态度量表、UWES-17、Koopmans IWPQ;68 个条目,12 个子量表),并在五级角色披露阶梯、呈现方式与推理努力消融、反事实人口学交换(性别、角色、教育)、跨语言检查以及逐字回忆记忆探测条件下,让涵盖 OpenAI、Anthropic、Google 及十二个开放权重系列共 37 个模型的阵容基于真实受访者画像进行条件生成。由此得到的心理测量相似性得分(PSS)以五个非 LLM 统计基线及一个留出的人类-人类比较上限为锚点,并提供了受访者自助法置信区间以及用于 Tucker's phi 的项目置换零分布。LLMs 能复现人类心理测量关系的定性方向,但在样本驱动的 PSS 组成部分上,高斯 copula 基线优于所有 LLM;LLM“群体”彼此之间的相似度(平均 LLM 间 PSS 为 0.73)高于其与人类的相似度;记忆化并未驱动排行榜(recall-PSS 秩相关为 0.00)。反事实交换揭示了教育驱动的效应(平均 |d|=0.56),远超性别(0.12)和角色(0.18);在 37 个模型中,有 8 个模型的 UWES Tucker's phi 落在置换零分布范围内。在下游分析中,每个 LLM 都表现出强烈的默许偏差(+0.84 个标准差);在合成数据上训练的回归模型在留出的人类样本上失去预测效度(平均 R² 为 -0.18,而人类为 0.28);并且模型在 10 条安慰剂中介路径中的 3 条上虚构出间接效应。LLM 样本并不能直接替代人类调查数据。
English
Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-psychology dataset (n=263 employees; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68 items, 12 subscales) and condition a 37-model lineup spanning OpenAI, Anthropic, Google, and twelve open-weight families on real respondent profiles under a five-level persona-disclosure ladder, presentation and reasoning-effort ablations, counterfactual demographic swaps (gender, role, education), a cross-language check, and a verbatim-recall memorization probe. The resulting Psychometric Similarity Score (PSS) is anchored against five non-LLM statistical baselines and a held-out human-vs-human ceiling, with respondent-bootstrap confidence intervals and an item-permutation null for Tucker's phi. LLMs reproduce the qualitative direction of human psychometric relationships, but a Gaussian-copula baseline beats every LLM on the sample-driven PSS components; the LLM "crowd" is more similar to itself (mean inter-LLM PSS 0.73) than to humans; and memorization does not drive the leaderboard (recall-PSS rank correlation 0.00). Counterfactual swaps reveal education-driven effects (mean |d|=0.56) that dwarf gender (0.12) and role (0.18); Tucker's phi on UWES falls inside the permutation null for 8 of 37 models. Downstream, every LLM shows a strong acquiescence shift (+0.84 SD), synthetic-trained regressors lose predictive validity on held-out humans (mean R^2 -0.18 vs 0.28), and models fabricate indirect effects on 3 of 10 placebo mediation paths. LLM samples are not a drop-in replacement for human survey data.