그럴듯하지만 타당하지 않은: LLM의 가상 설문 응답자 활용에 대한 심리측정학적 감사
Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents
July 6, 2026
저자: Mantas Lukauskas, Viktorija Šarkauskaitė
cs.AI
초록
대규모 언어 모델(LLM)은 점점 더 합성 설문 응답자로 활용되고 있지만, 기존 평가는 개인 수준에서 응답이 그럴듯해 보이는지 여부만 묻는다. 우리는 올바른 질문은 심리측정적(psychometric)이어야 한다고 주장한다. 즉, LLM이 실제 인간 설문 데이터의 결합 분포, 잠재 구조, 신뢰도, 매개 경로, 인구통계학적 효과를 보존하는가? 우리는 리투아니아 조직심리학 데이터셋(n=263명의 직원; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68개 문항, 12개 하위 척도)을 도입하고, OpenAI, Anthropic, Google 및 12개 오픈가중치 계열에 속하는 37개 모델 라인업을 실제 응답자 프로필에 조건화하여, 5단계 페르소나 공개 사다리, 제시 및 추론 노력 제거(ablation) 실험, 반사실적 인구통계 교체(성별, 직무, 교육), 교차 언어 검증, 축어적 회상 암기 탐지 절차를 적용하였다. 그 결과 산출된 심리측정 유사도 점수(PSS)는 5개의 비(非)LLM 통계적 기준선과 홀드아웃 인간-인간 상한선에 기준을 두고, 응답자 부트스트랩 신뢰 구간과 Tucker's phi에 대한 항목 순열 귀무 분포를 사용하여 검증되었다. LLM은 인간의 심리측정 관계에서 질적 방향성은 재현하지만, 가우스 코퓰라 기준선이 표본 기반 PSS 구성 요소에서 모든 LLM을 능가하였다. LLM '집단'은 인간보다 자기 자신과 더 유사하며(LLM 간 평균 PSS 0.73), 암기는 리더보드를 주도하지 않았다(회상-PSS 순위 상관 0.00). 반사실적 교체에서 교육 수준에 따른 효과(평균 |d|=0.56)가 성별(0.12)과 직무(0.18)을 압도하였고, UWES에 대한 Tucker's phi는 37개 모델 중 8개에서 순열 귀무 분포 내에 있었다. 하위 분석에서 모든 LLM은 강한 동의 편향 이동(+0.84 SD)을 보였으며, 합성 데이터로 훈련된 회귀 모델은 홀드아웃 인간 데이터에 대한 예측 타당도를 상실하였고(평균 R^2 -0.18 대 0.28), 10개의 위약 매개 경로 중 3개에서 허위 간접 효과를 생성하였다. LLM 표본은 인간 설문 데이터의 대체 투입물이 될 수 없다.
English
Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-psychology dataset (n=263 employees; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68 items, 12 subscales) and condition a 37-model lineup spanning OpenAI, Anthropic, Google, and twelve open-weight families on real respondent profiles under a five-level persona-disclosure ladder, presentation and reasoning-effort ablations, counterfactual demographic swaps (gender, role, education), a cross-language check, and a verbatim-recall memorization probe. The resulting Psychometric Similarity Score (PSS) is anchored against five non-LLM statistical baselines and a held-out human-vs-human ceiling, with respondent-bootstrap confidence intervals and an item-permutation null for Tucker's phi. LLMs reproduce the qualitative direction of human psychometric relationships, but a Gaussian-copula baseline beats every LLM on the sample-driven PSS components; the LLM "crowd" is more similar to itself (mean inter-LLM PSS 0.73) than to humans; and memorization does not drive the leaderboard (recall-PSS rank correlation 0.00). Counterfactual swaps reveal education-driven effects (mean |d|=0.56) that dwarf gender (0.12) and role (0.18); Tucker's phi on UWES falls inside the permutation null for 8 of 37 models. Downstream, every LLM shows a strong acquiescence shift (+0.84 SD), synthetic-trained regressors lose predictive validity on held-out humans (mean R^2 -0.18 vs 0.28), and models fabricate indirect effects on 3 of 10 placebo mediation paths. LLM samples are not a drop-in replacement for human survey data.