EpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections

📅 2026-07-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the frequent issues of factual inconsistency and information omission in narratives generated by large language models for epidemiological forecasting. To mitigate these problems, the authors propose an agent-based framework that decouples reasoning from text generation. The approach first extracts scenario dimensions from multidimensional forecast data and constructs a partial order structure to systematically traverse the scenario space. It then employs a semantic-arithmetic consistency-aware grammar to produce quantitative statements and selects high-information content via a maximum entropy principle. Key innovations include partial-order scenario modeling, consistency-aware comparative grammar, and entropy-driven content selection. Experiments on data from the COVID-19 Scenario Modeling Hub demonstrate that the generated narratives significantly outperform baseline methods in factual accuracy and coverage of critical patterns, while preserving the stylistic qualities of expert reports.
📝 Abstract
Generation of clear and accessible public health narratives is critical for communicating complex epidemiological projections to policymakers and the general public at large. Such narratives require more than simply reporting numbers: projections must be contextualized and quantitatively grounded across multiple dimensions. Further, projections are often derived from large ensemble datasets which combine intervention assumptions, geographic and demographic strata, outcomes, time horizons, and uncertainty quantiles. However, directly using large language models (LLMs) to summarize and contextualize such data often leads to inconsistencies, omissions, and fragile behavior. We introduce an agentic framework (EpiNarrate) for public health report generation that separates structured numerical reasoning from natural-language generation. The framework first extracts scenario axes and organizes them into a partial-order schema, enabling systematic traversal of the underlying multidimensional space. It then constructs an augmented dataset and derives valid quantitative statements through a comparison grammar that enforces semantic and arithmetic consistency. To balance coverage and non-redundancy, we introduce an interestingness-driven selection mechanism based on maximum-entropy principles. Experiments on the COVID-19 Scenario Modeling Hub demonstrate that our model produces narratives with improved factual grounding and broader coverage of salient epidemiological patterns, while preserving the style of expert-written reports.
Problem

Research questions and friction points this paper is trying to address.

epidemiological projections
public health narratives
large language models
multidimensional data
factual grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

agentic framework
grounded narrative generation
epidemiological projections
structured reasoning
maximum-entropy selection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Rituparna Datta
Rituparna Datta
University of South Alabama
Engineering OptimizationEvolutionary ComputationNeural NetworksManufacturingRobotics
S
Srini Venkatramanan
Biocomplexity Institute, University of Virginia
B
Bryan L. Lewis
Biocomplexity Institute, University of Virginia
Y
Yiqi Su
Department of Computer Science, Virginia Tech
Harry Hochheiser
Harry Hochheiser
Professor of Biomedical Informatics, University of Pittsburgh
Biomedical informaticsbioinformaticsclinical informaticshuman-computer interaction
L
Lucie Contamin
Department of Biomedical Informatics, University of Pittsburgh
P
Parantapa Bhattacharya
Biocomplexity Institute, University of Virginia
Naren Ramakrishnan
Naren Ramakrishnan
Thomas L. Phillips Professor, Virginia Tech
ForecastingMachine LearningComputational epidemiologyRecommender systemsVisual analytics
A
Anil Vullikanti
Biocomplexity Institute and Department of Computer Science, University of Virginia