Generating Public Health Responses using Survey-Augmented Large Language Models

📅 2026-06-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Traditional large-scale health surveys are costly and time-consuming, rendering them ill-suited for rapidly evolving epidemic scenarios. This study proposes a novel approach that integrates cluster analysis with prompt engineering to guide large language models in generating synthetic public health survey data exhibiting population-level behavioral consistency, particularly in domains such as vaccine attitudes, risk perception, and health-related behaviors. The method uniquely incorporates a clustering-informed mechanism into prompt design and is validated using real longitudinal influenza survey data. Experimental results demonstrate that the generated data effectively replicates demographic characteristics and univariate distributions of individual variables, with certain models capturing aggregate vaccination trends. However, limitations remain in modeling joint relationships among variables, and the synthetic nature of the data remains detectable by classifiers.
📝 Abstract
Epidemiological models often rely on survey data to represent how individuals make health-related decisions, such as whether to vaccinate or adopt protective behaviors. However, repeated large-scale surveys are costly, time-consuming, and limited in the range of scenarios they can capture. In this work, we investigate whether large language models (LLMs) can generate synthetic survey responses that reproduce patterns observed in real populations. Using longitudinal data from the FluPaths surveys, we first identify groups associated with broadly positive or negative attitudes toward vaccination through clustering analysis. We then evaluate several LLMs using a cluster-informed prompting approach to generate synthetic survey responses across multiple epidemic waves. Across models, the synthetic data generally reproduce the distributions of demographic characteristics, vaccination-related beliefs, risk perceptions, and health behaviors observed in the survey data. However, they are less successful at capturing how these factors vary together within respondents. Some models reproduce group-level vaccination trends more reliably than others, although performance varies across waves. We also trained a classifier to distinguish real from synthetic records and found that the generated responses remained identifiable as synthetic. Overall, our findings suggest that LLM-generated survey data may provide a useful tool for exploratory data augmentation and we hope that it could support agent-based epidemic modeling approaches. However, the generated data should not be treated as a substitute for human survey data without further methodological improvements and validation.
Problem

Research questions and friction points this paper is trying to address.

survey data
large language models
synthetic responses
epidemiological modeling
health behaviors
Innovation

Methods, ideas, or system contributions that make the work stand out.

survey-augmented LLMs
synthetic survey data
cluster-informed prompting
epidemic modeling
vaccination attitudes
💼 Related Jobs
No related jobs found.
L
Leonardo Marciaga
Illinois Institute of Technology
T
Thuyen Pham
University of Massachusetts Amherst
J
Julia Rezvani
Portland State University
A
Alina Hyk
Oregon State University
Chunyang Liao
Chunyang Liao
UCLA
Approximation TheoryMathematical Data Science
K
Konstantinos Mitsopoulos
Johns Hopkins University
R
Raffaele Vardavas
Causal Paths Analytics LLC