ChatPaper.aiChatPaper

看似合理但缺乏效度:以LLM作為合成調查受訪者之心理計量審計

Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents

July 6, 2026
作者: Mantas Lukauskas, Viktorija Šarkauskaitė
cs.AI

摘要

大型語言模型(LLMs)日益被用作合成調查受訪者,但現有的評估僅在個體層面檢視答案是否看似合理。我們主張正確的問題應是心理測量學層面的:LLMs 是否保留了真實人類調查資料的聯合分布、潛在結構、信度、中介路徑及人口統計效應?我們引入一個立陶宛組織心理學資料集(n=263 名員工;包含 Dunham 變革態度量表、UWES-17、Koopmans IWPQ;共 68 題、12 個分量表),並在五級人物揭露階梯、呈現方式與推理努力度的消融實驗、反事實人口統計交換(性別、角色、教育程度)、跨語言檢核,以及逐字回憶記憶探測等條件下,讓橫跨 OpenAI、Anthropic、Google 及十二個開放權重系列的 37 個模型陣容基於真實受訪者輪廓進行作答。由此產生的心理測量相似度分數(PSS)以五個非 LLM 統計基線及一個保留的人類對人類上限為錨點,並使用受訪者自助抽樣法計算信賴區間,以及針對 Tucker's phi 的項目排列虛無假設檢定。LLMs 能再現人類心理測量關係的定性方向,但高斯連結函數基線在樣本驅動的 PSS 成分上擊敗所有 LLM;LLM「群體」彼此之間的相似度(平均 LLM 間 PSS 為 0.73)高於其與人類的相似度;而記憶效應並非驅動排行榜的因素(回憶-PSS 等級相關係數為 0.00)。反事實交換顯示教育程度驅動的效應(平均 |d|=0.56)遠大於性別(0.12)與角色(0.18);在 UWES 上,37 個模型中有 8 個的 Tucker's phi 落在排列虛無假設的區間內。在下游應用中,每個 LLM 都表現出強烈的默許偏差(+0.84 個標準差),以合成資料訓練的迴歸模型在保留的人類樣本上喪失預測效度(平均 R^2 從 0.28 降至 -0.18),且模型在 10 條安慰劑中介路徑中的 3 條上捏造出間接效應。LLM 樣本不能直接替代人類調查資料。
English
Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-psychology dataset (n=263 employees; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68 items, 12 subscales) and condition a 37-model lineup spanning OpenAI, Anthropic, Google, and twelve open-weight families on real respondent profiles under a five-level persona-disclosure ladder, presentation and reasoning-effort ablations, counterfactual demographic swaps (gender, role, education), a cross-language check, and a verbatim-recall memorization probe. The resulting Psychometric Similarity Score (PSS) is anchored against five non-LLM statistical baselines and a held-out human-vs-human ceiling, with respondent-bootstrap confidence intervals and an item-permutation null for Tucker's phi. LLMs reproduce the qualitative direction of human psychometric relationships, but a Gaussian-copula baseline beats every LLM on the sample-driven PSS components; the LLM "crowd" is more similar to itself (mean inter-LLM PSS 0.73) than to humans; and memorization does not drive the leaderboard (recall-PSS rank correlation 0.00). Counterfactual swaps reveal education-driven effects (mean |d|=0.56) that dwarf gender (0.12) and role (0.18); Tucker's phi on UWES falls inside the permutation null for 8 of 37 models. Downstream, every LLM shows a strong acquiescence shift (+0.84 SD), synthetic-trained regressors lose predictive validity on held-out humans (mean R^2 -0.18 vs 0.28), and models fabricate indirect effects on 3 of 10 placebo mediation paths. LLM samples are not a drop-in replacement for human survey data.