ChatPaper.aiChatPaper

もっともらしいが妥当ではない:合成調査回答者としてのLLMの心理測定的監査

Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents

July 6, 2026
著者: Mantas Lukauskas, Viktorija Šarkauskaitė
cs.AI

要旨

大規模言語モデル(LLM)は合成的な調査回答者としてますます用いられているが、既存の評価は回答が個別レベルで妥当に見えるかどうかを問うにとどまる。我々は、正しい問いは心理測定的なものであると主張する。すなわち、LLMは実際の人間の調査データの同時分布、潜在構造、信頼性、媒介経路、および人口統計学的効果を保持しているだろうか。我々は、リトアニアの組織心理学データセット(従業員263名;Dunham Attitudes Toward Change、UWES-17、Koopmans IWPQ;68項目、12下位尺度)を導入し、OpenAI、Anthropic、Google、および12のオープンウェイトファミリーにわたる37モデルのラインナップを、5段階のペルソナ開示条件、提示および推論努力のアブレーション、反事実的人口統計学的交換(性別、役割、教育)、言語横断チェック、および逐語的想起記憶プローブの下で、実際の回答者プロファイルに条件付けた。得られた心理測定的類似度スコア(PSS)は、5つの非LLM統計ベースラインと、ホールドアウトされた人間対人間の上限に対して基準化され、回答者ブートストラップ信頼区間と、タッカーのφのための項目並べ替え帰無仮説を伴う。LLMは人間の心理測定的関係の質的な方向を再現するが、サンプル駆動のPSS構成要素ではガウス・コピュラベースラインがすべてのLLMを上回る。LLMの「群衆」は人間よりもそれ自体に類似しており(LLM間平均PSS 0.73)、記憶化はリーダーボードを駆動しない(想起-PSS順位相関0.00)。反事実的交換は、性別(0.12)と役割(0.18)を凌駕する教育駆動効果(平均|d|=0.56)を明らかにする。UWESのタッカーのφは、37モデル中8モデルで並べ替え帰無仮説の範囲内に収まる。下流では、すべてのLLMが強い同調傾向のシフト(+0.84 SD)を示し、合成データで訓練された回帰モデルはホールドアウトされた人間データに対する予測妥当性を失い(平均R^2 -0.18、対0.28)、またモデルは10のプラセボ媒介経路のうち3つで間接効果を捏造する。LLMサンプルは、人間の調査データのそのままの代替にはならない。
English
Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-psychology dataset (n=263 employees; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68 items, 12 subscales) and condition a 37-model lineup spanning OpenAI, Anthropic, Google, and twelve open-weight families on real respondent profiles under a five-level persona-disclosure ladder, presentation and reasoning-effort ablations, counterfactual demographic swaps (gender, role, education), a cross-language check, and a verbatim-recall memorization probe. The resulting Psychometric Similarity Score (PSS) is anchored against five non-LLM statistical baselines and a held-out human-vs-human ceiling, with respondent-bootstrap confidence intervals and an item-permutation null for Tucker's phi. LLMs reproduce the qualitative direction of human psychometric relationships, but a Gaussian-copula baseline beats every LLM on the sample-driven PSS components; the LLM "crowd" is more similar to itself (mean inter-LLM PSS 0.73) than to humans; and memorization does not drive the leaderboard (recall-PSS rank correlation 0.00). Counterfactual swaps reveal education-driven effects (mean |d|=0.56) that dwarf gender (0.12) and role (0.18); Tucker's phi on UWES falls inside the permutation null for 8 of 37 models. Downstream, every LLM shows a strong acquiescence shift (+0.84 SD), synthetic-trained regressors lose predictive validity on held-out humans (mean R^2 -0.18 vs 0.28), and models fabricate indirect effects on 3 of 10 placebo mediation paths. LLM samples are not a drop-in replacement for human survey data.