Mapping the Emerging Social Science of Large Language Models

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过多种分析方法,对大型语言模型的社会科学研究进行分类和映射,识别出三大研究领域及其子类别,以理解LLM的社会影响。
📝 Abstract
Large language models (LLMs) increasingly shape communication, learning, work, creativity, and decision-making, yet social-science research on these developments remains fragmented. We map this emerging field using a curated corpus of 198 papers reviewed in full and a field-scale corpus of 47,719 published papers from five bibliographic databases. Combining sentence embeddings, K-means clustering, within-cluster Latent Dirichlet Allocation (LDA), author and LLM classifications, and structural topic modeling, we identify three domains: LLM as Social Minds, examining socially interpretable model behavior; LLM Societies, examining collective dynamics among interacting model-based agents; and LLM-Human Interactions, examining how people perceive, use, and are affected by LLMs. These domains contain 13 subcategories spanning reasoning, personality and bias, behavioral games, collective intelligence, simulation, trust, work, creativity, and education. In the curated corpus, the three-domain solution is highly stable under resampling (adjusted Rand index = 0.952), and K-means assignments agree with author full-text classifications for 77.78% of papers. At field scale, 13 of 15 topics map onto the taxonomy, while K-means and structural-topic-model domains agree for 73.83% of overlapping papers. LLM-Human Interactions accounts for 78.02% of domain-mapped topic mass, but venue analysis reveals a contrasting pattern: Social Minds and LLM Societies together account for 66.37% of highly cited papers in leading conference venues, whereas LLM-Human Interactions accounts for 76.81% in the corresponding journal subset. The resulting taxonomy provides a reproducible framework for understanding how model behavior, agent interaction, and institutional context jointly shape the social consequences of LLMs.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Social Science Research
Fragmented Studies
Systematic Framework
Social Consequences
Innovation

Methods, ideas, or system contributions that make the work stand out.

sentence embeddings
K-means clustering
Latent Dirichlet Allocation (LDA)
structural topic modeling
🔎 Similar Papers
No similar papers found.
Y
Yi Yang
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, Shenzhen, China
Xiao Jia
Xiao Jia
Stanford University
deep learningmedical image analysis
Z
Zeyun Dong
School of Computer Science and Technology, Xidian University, Xi’an, China
C
Chenzhang Wang
School of Mathematics, The University of Edinburgh, Edinburgh, United Kingdom
Zhanzhan Zhao
Zhanzhan Zhao
Harvard Medical School
emergent computingsegregationcomputational social science