HealthLoopQA: A Context-Aware Question Answering Benchmark for Interpreting Wearable Monitoring Data in Diabetes Care

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决长期医疗监测数据解释问题,研究通过创建HealthLoopQA基准来评估大型语言模型在糖尿病护理中的推理能力。
📝 Abstract
As medical wearables become integrated into daily chronic disease care, effectively interpreting longitudinal monitoring data is essential for patients and clinicians to understand health trends, detect safety-critical events, and make informed decisions. While large language models (LLMs) show promise for transforming this streaming physiological data into personalized health insights, evaluating their reasoning capability and analytical rigor in diverse monitoring tasks remains a fundamental challenge. Existing medical wearable question answering (QA) benchmarks primarily assess short-horizon classification or statistical summaries, largely ignoring the long-term patterns, therapeutic and behavioural contexts, and potential system failures inherent in real-world deployments. To address this, we introduce HealthLoopQA, a comprehensive diagnostic benchmark for evaluating LLM reasoning over continuous diabetes monitoring data. Grounded in a novel taxonomy of eleven atomic reasoning abilities, HealthLoopQA comprises 127 tasks and over 1,500 QA instances spanning process mining, anomaly detection, and prediction over 30-day horizons. To systematically evaluate safety awareness, we complement real-world datasets with a fault-injected simulation testbed modeling diverse device malfunctions and cyber-physical attacks to generate physiologically plausible hazard scenarios. Evaluating state-of-the-art LLMs across prompting and agentic frameworks reveals severe limitations in complex temporal pattern mining. Furthermore, we identify a broader phenomenon of In-context Laziness under long-context prompting, highlighting critical open challenges in deploying LLMs for rigorous long-horizon medical reasoning.
Problem

Research questions and friction points this paper is trying to address.

medical wearables
longitudinal monitoring data
large language models
reasoning capability
diabetes care
Innovation

Methods, ideas, or system contributions that make the work stand out.

Context-Aware
Wearable Monitoring Data
Diabetes Care
Large Language Models
In-context Laziness
Y
Yuchen Niu
Nanyang Technological University
Yanan Ma
Yanan Ma
City University of Hong Kong
Wireless networksEdge intelligence
S
Srinivasan Nandakumar
Imperial College London, Imperial Global Singapore
M
Maolin Chen
National University of Singapore
Viktor Schlegel
Viktor Schlegel
Deputy Director IN-CYPHER Programme @ IGS, Imperial College London
Natural Language UnderstandingAI for HealthcareClinical NLPAI Evaluation
K
Kexin Wei
Imperial College London, Imperial Global Singapore
L
Ling Cheng
Imperial College London, Imperial Global Singapore
A
Anna Bird
Imperial College London, Imperial Global Singapore
A
Anil Anthony Bharath
Imperial College London, Imperial Global Singapore
Siew-Kei Lam
Siew-Kei Lam
Nanyang Technological University
Custom ComputingEmbedded VisionEdge AIEmbedded System SecurityTransportation Analytics