Hugging Face Daily Papers · · 4 min read

Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

LLMs can generate very plausible survey responses - but are they psychometrically valid? We find that plausibility often hides unstable constructs, weak reliability, and poor agreement with real human response patterns.<br>The paper was originally submitted to arXiv in July, but remained on moderation hold and was only publicly released today.</p>\n","updatedAt":"2026-08-18T13:42:20.400Z","author":{"_id":"636f4e01af79dd1936034ef7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/6VFfzcEbMowp0scR7Lx7J.png","fullname":"Mantas Lukauskas","name":"MLukauskas","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65d8469c8f6d9b89466e6478/ihMnOK12kGcuhFC3nNWf2.png","fullname":"Hostinger","name":"Hostinger","type":"org","isHf":false,"plan":"enterprise"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9838992953300476},"editors":["MLukauskas"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/6VFfzcEbMowp0scR7Lx7J.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.14606","authors":[{"_id":"6a83ee9c675db694db8cd601","user":{"_id":"636f4e01af79dd1936034ef7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/6VFfzcEbMowp0scR7Lx7J.png","isPro":false,"fullname":"Mantas Lukauskas","user":"MLukauskas","type":"user","name":"MLukauskas"},"name":"Mantas Lukauskas","status":"claimed_verified","statusLastChangedAt":"2026-08-18T08:45:05.204Z","hidden":false},{"_id":"6a83ee9c675db694db8cd602","name":"Viktorija Šarkauskaitė","hidden":false}],"publishedAt":"2026-07-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents","submittedOnDailyBy":{"_id":"636f4e01af79dd1936034ef7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/6VFfzcEbMowp0scR7Lx7J.png","isPro":false,"fullname":"Mantas Lukauskas","user":"MLukauskas","type":"user","name":"MLukauskas"},"summary":"Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-psychology dataset (n=263 employees; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68 items, 12 subscales) and condition a 37-model lineup spanning OpenAI, Anthropic, Google, and twelve open-weight families on real respondent profiles under a five-level persona-disclosure ladder, presentation and reasoning-effort ablations, counterfactual demographic swaps (gender, role, education), a cross-language check, and a verbatim-recall memorization probe. The resulting Psychometric Similarity Score (PSS) is anchored against five non-LLM statistical baselines and a held-out human-vs-human ceiling, with respondent-bootstrap confidence intervals and an item-permutation null for Tucker's phi. LLMs reproduce the qualitative direction of human psychometric relationships, but a Gaussian-copula baseline beats every LLM on the sample-driven PSS components; the LLM \"crowd\" is more similar to itself (mean inter-LLM PSS 0.73) than to humans; and memorization does not drive the leaderboard (recall-PSS rank correlation 0.00). Counterfactual swaps reveal education-driven effects (mean |d|=0.56) that dwarf gender (0.12) and role (0.18); Tucker's phi on UWES falls inside the permutation null for 8 of 37 models. Downstream, every LLM shows a strong acquiescence shift (+0.84 SD), synthetic-trained regressors lose predictive validity on held-out humans (mean R^2 -0.18 vs 0.28), and models fabricate indirect effects on 3 of 10 placebo mediation paths. LLM samples are not a drop-in replacement for human survey data.","upvotes":1,"discussionId":"6a83ee9c675db694db8cd603","ai_summary":"Large language models replicate broad psychometric trends in synthetic survey responses but fail to match human joint distributions, reliability, and mediation structures, making them unsuitable replacements for real respondents.","ai_keywords":["large language models","psychometric similarity score","Gaussian copula","Tucker's phi","acquiescence shift","mediation pathways","counterfactual demographic swaps","synthetic survey respondents"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"636f4e01af79dd1936034ef7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/6VFfzcEbMowp0scR7Lx7J.png","isPro":false,"fullname":"Mantas Lukauskas","user":"MLukauskas","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.14606.md","query":{}}">
Papers
arxiv:2608.14606

Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents

Published on Jul 6
· Submitted by
Mantas Lukauskas
on Aug 18
Authors:

Abstract

Large language models replicate broad psychometric trends in synthetic survey responses but fail to match human joint distributions, reliability, and mediation structures, making them unsuitable replacements for real respondents.

Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-psychology dataset (n=263 employees; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68 items, 12 subscales) and condition a 37-model lineup spanning OpenAI, Anthropic, Google, and twelve open-weight families on real respondent profiles under a five-level persona-disclosure ladder, presentation and reasoning-effort ablations, counterfactual demographic swaps (gender, role, education), a cross-language check, and a verbatim-recall memorization probe. The resulting Psychometric Similarity Score (PSS) is anchored against five non-LLM statistical baselines and a held-out human-vs-human ceiling, with respondent-bootstrap confidence intervals and an item-permutation null for Tucker's phi. LLMs reproduce the qualitative direction of human psychometric relationships, but a Gaussian-copula baseline beats every LLM on the sample-driven PSS components; the LLM "crowd" is more similar to itself (mean inter-LLM PSS 0.73) than to humans; and memorization does not drive the leaderboard (recall-PSS rank correlation 0.00). Counterfactual swaps reveal education-driven effects (mean |d|=0.56) that dwarf gender (0.12) and role (0.18); Tucker's phi on UWES falls inside the permutation null for 8 of 37 models. Downstream, every LLM shows a strong acquiescence shift (+0.84 SD), synthetic-trained regressors lose predictive validity on held-out humans (mean R^2 -0.18 vs 0.28), and models fabricate indirect effects on 3 of 10 placebo mediation paths. LLM samples are not a drop-in replacement for human survey data.

Community

Paper author Paper submitter about 1 hour ago

LLMs can generate very plausible survey responses - but are they psychometrically valid? We find that plausibility often hides unstable constructs, weak reliability, and poor agreement with real human response patterns.
The paper was originally submitted to arXiv in July, but remained on moderation hold and was only publicly released today.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.14606
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.14606 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.14606 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.14606 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers