🤖 AI Summary
This study addresses the reconstruction of collective voting behavior using large language models (LLMs) by proposing a methodological framework that treats LLMs as compressed representations of social reality. Through demographic-conditional generation, probabilistic preference elicitation, and soft voting aggregation, the approach infers aggregate political tendencies from individual profiles. Applied to Czech elections, the method replicates official outcomes with minimal error, accurately recovering political bloc structures and sociodemographic gradients. These findings validate the efficacy of LLMs as latent sociological models, offering computational social science a novel instrument for exploring complex social systems.
📝 Abstract
Large language models (LLMs) trained on large-scale internet corpora encode extensive statistical regularities about social identities, attitudes, and political behaviour. This paper introduces and evaluates a methodological framework that leverages these latent representations to reconstruct aggregate voting behaviour from individual-level sociodemographic profiles. We operationalize LLMs as implicit sociological models by conditioning them on demographic descriptions, eliciting probabilistic turnout and party preferences, and aggregating individual outputs via a soft voting procedure. Using the 2021 Czech parliamentary election as a validation case, we demonstrate that contemporary LLMs reproduce official election outcomes with low mean absolute error, recover known political bloc structures, and align with independently established sociodemographic gradients. The contribution of this work is methodological rather than predictive: we show how LLMs can be systematically interrogated as compressed representations of social reality, offering a novel exploratory instrument for computational social science while clearly delineating its epistemic and ethical limits.