Large language models simulate intersectional synthetic identities with a budget of one to two dimensions

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究测试了大语言模型在模拟交叉群体意见时的表现,发现模型仅能有效表示单一身份特征,无法准确反映多维度交叉身份的意见特性。
📝 Abstract
Large language models are increasingly used as synthetic survey respondents, promising cheap access to rare intersectional populations. We test standard demographic-persona methods against every real intersectional subgroup across 15 waves of Pew's American Trends Panel -- 21 million simulated response distributions from eight models. In real respondents, subgroup opinion is approximately the additive sum of its single-identity components, yet grows 2.5x more distinctive as identities intersect. Simulated respondents show no such composition: a single feature explains a two-feature persona's responses better than the additive combination in 75-82% of subgroups, and a third feature adds almost nothing. This collapse survives every prompting strategy we test. Additionally, the feature models retain is chosen nearly blindly -- except that they systematically discard race and religion, the strongest real drivers of opinion. Synthetic samples offer intersectional personas but represent one identity at a time.
Problem

Research questions and friction points this paper is trying to address.

large language models
intersectional identities
synthetic respondents
opinion distribution
demographic-persona methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

large language models
intersectional identities
synthetic respondents
opinion formation
demographic-persona methods