🤖 AI Summary
Traditional large-scale health surveys are costly and time-consuming, rendering them ill-suited for rapidly evolving epidemic scenarios. This study proposes a novel approach that integrates cluster analysis with prompt engineering to guide large language models in generating synthetic public health survey data exhibiting population-level behavioral consistency, particularly in domains such as vaccine attitudes, risk perception, and health-related behaviors. The method uniquely incorporates a clustering-informed mechanism into prompt design and is validated using real longitudinal influenza survey data. Experimental results demonstrate that the generated data effectively replicates demographic characteristics and univariate distributions of individual variables, with certain models capturing aggregate vaccination trends. However, limitations remain in modeling joint relationships among variables, and the synthetic nature of the data remains detectable by classifiers.
📝 Abstract
Epidemiological models often rely on survey data to represent how individuals make health-related decisions, such as whether to vaccinate or adopt protective behaviors. However, repeated large-scale surveys are costly, time-consuming, and limited in the range of scenarios they can capture. In this work, we investigate whether large language models (LLMs) can generate synthetic survey responses that reproduce patterns observed in real populations. Using longitudinal data from the FluPaths surveys, we first identify groups associated with broadly positive or negative attitudes toward vaccination through clustering analysis. We then evaluate several LLMs using a cluster-informed prompting approach to generate synthetic survey responses across multiple epidemic waves. Across models, the synthetic data generally reproduce the distributions of demographic characteristics, vaccination-related beliefs, risk perceptions, and health behaviors observed in the survey data. However, they are less successful at capturing how these factors vary together within respondents. Some models reproduce group-level vaccination trends more reliably than others, although performance varies across waves. We also trained a classifier to distinguish real from synthetic records and found that the generated responses remained identifiable as synthetic. Overall, our findings suggest that LLM-generated survey data may provide a useful tool for exploratory data augmentation and we hope that it could support agent-based epidemic modeling approaches. However, the generated data should not be treated as a substitute for human survey data without further methodological improvements and validation.