Institution profile

UK Health Security Agency

Academic institutioneurope · gb
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information

May 09, 2025

This study addresses the limited knowledge acquisition and trustworthy generation capabilities of large language models (LLMs) in UK public health. We introduce PubHealthBench—the first domain-specific benchmark for UK government health policy—comprising 8,000+ high-quality question-answer pairs. Methodologically, we propose an automated question-generation pipeline and a dual-mode evaluation framework integrating multiple-choice question answering (MCQA) and open-ended generation. Our approach leverages official guidance documents for structured knowledge extraction, annotation, and open-sourcing of the underlying policy corpus. Results show that state-of-the-art closed-source models (e.g., GPT-4.5) achieve >90% accuracy on MCQA—outperforming generic search engines—but score <75% on open-ended generation, exposing critical gaps in factual consistency and authoritative grounding. Our key contribution is the first automated, policy-grounded evaluation paradigm for UK public health, empirically revealing a fundamental disparity between discriminative and generative capabilities—thereby establishing a foundational benchmark and actionable pathway toward trustworthy health AI.

0 citationsRead paper

Review GIDE -- Restaurant Review Gastrointestinal Illness Detection and Extraction with Large Language Models

Mar 12, 2025

This study addresses the challenge of underreported and undiagnosed foodborne gastrointestinal (GI) illnesses—cases often missed by conventional public health surveillance. We propose a passive disease signal detection paradigm leveraging online restaurant reviews and large language models (LLMs). Methodologically, we introduce the first multi-level expert annotation schema tailored for GI diseases, enabling joint extraction of disease mentions, symptom expressions, and implicated foods. Using open-source LLMs (e.g., Llama, Phi), we implement end-to-end information extraction via zero-shot and few-shot prompting, and rigorously evaluate model robustness across gender, geographic, and cuisine-related biases. Experimental results show micro-F1 scores exceeding 90% across all three tasks, with prompting outperforming fine-tuned RoBERTa baselines. Our key contribution is demonstrating that lightweight prompting strategies achieve high accuracy and strong generalizability in fine-grained health information extraction—establishing a scalable, low-cost technical foundation for broad-coverage digital epidemiological surveillance.

0 citationsRead paper
Recent publications

Latest Papers

Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information

May 09, 2025

This study addresses the limited knowledge acquisition and trustworthy generation capabilities of large language models (LLMs) in UK public health. We introduce PubHealthBench—the first domain-specific benchmark for UK government health policy—comprising 8,000+ high-quality question-answer pairs. Methodologically, we propose an automated question-generation pipeline and a dual-mode evaluation framework integrating multiple-choice question answering (MCQA) and open-ended generation. Our approach leverages official guidance documents for structured knowledge extraction, annotation, and open-sourcing of the underlying policy corpus. Results show that state-of-the-art closed-source models (e.g., GPT-4.5) achieve >90% accuracy on MCQA—outperforming generic search engines—but score <75% on open-ended generation, exposing critical gaps in factual consistency and authoritative grounding. Our key contribution is the first automated, policy-grounded evaluation paradigm for UK public health, empirically revealing a fundamental disparity between discriminative and generative capabilities—thereby establishing a foundational benchmark and actionable pathway toward trustworthy health AI.

0 citationsRead paper

Review GIDE -- Restaurant Review Gastrointestinal Illness Detection and Extraction with Large Language Models

Mar 12, 2025

This study addresses the challenge of underreported and undiagnosed foodborne gastrointestinal (GI) illnesses—cases often missed by conventional public health surveillance. We propose a passive disease signal detection paradigm leveraging online restaurant reviews and large language models (LLMs). Methodologically, we introduce the first multi-level expert annotation schema tailored for GI diseases, enabling joint extraction of disease mentions, symptom expressions, and implicated foods. Using open-source LLMs (e.g., Llama, Phi), we implement end-to-end information extraction via zero-shot and few-shot prompting, and rigorously evaluate model robustness across gender, geographic, and cuisine-related biases. Experimental results show micro-F1 scores exceeding 90% across all three tasks, with prompting outperforming fine-tuned RoBERTa baselines. Our key contribution is demonstrating that lightweight prompting strategies achieve high accuracy and strong generalizability in fine-grained health information extraction—establishing a scalable, low-cost technical foundation for broad-coverage digital epidemiological surveillance.

0 citationsRead paper