FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建FairGlucose基准,评估了33种CGM模型在不同人群中的公平性,揭示了仅凭总体水平验证无法发现的子群体差异问题。
📝 Abstract
As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To enable this evaluation, we constructed FairGlucose, a 300-patient CGM cohort balanced across 12 demographic strata (age x gender x type 1/type 2 diabetes), with 132,480 forecasting samples and 3,945 unique behavioral events (meals, exercise, medication) logged by 81 patients. Benchmarking 33 models across four families on 2-hour glucose forecasting, we find that population-level external validation can conceal substantial subgroup disparities. Aggregate out-of-distribution metrics appear stable (approximately 1.0), yet subgroup-level ratios range from 0.8 to 1.4, with T1D patients showing 6 mg/dL higher prediction error than T2D (p < 0.001). This disparity persists across all 33 models, suggesting a property of the prediction task rather than any single architecture. Further analysis shows that subgroup performance gaps align with the proportion of clinically hard cases, and that input-length sensitivity varies across demographics, motivating personalized configurations. Frontier LLMs underperform specialized neural models by 1-6 mg/dL; behavioral events contribute negligibly (approximately 0.1 mg/dL) even under oracle event access. These findings establish that population-level validation alone is insufficient for equity assessment of digital health AI, motivating subgroup-disaggregated reporting as a default standard.
Problem

Research questions and friction points this paper is trying to address.

CGM
AI tools
demographic disparities
accuracy equity
subgroup analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

CGM Fairness Benchmark
Subgroup Disparities
Population-Level Validation
Personalized Configurations
Behavioral Events
🔎 Similar Papers
No similar papers found.
Junjie Luo
Junjie Luo
Zhejiang A&F University
Urban digital twinVisual perceptionUAV modelling
X
Xuzhe Zhi
Carey Business School, Johns Hopkins University, Baltimore, MD, USA
R
Rui Han
Welldoc, Inc., Columbia, MD, USA
A
Abhimanyu Kumbara
Welldoc, Inc., Columbia, MD, USA
A
Anand K. Iyer
Welldoc, Inc., Columbia, MD, USA
M
Mansur E. Shomali
Welldoc, Inc., Columbia, MD, USA
Ritu Agarwal
Ritu Agarwal
Johns Hopkins Carey Business School
Healthcare information technologyhealth analyticsAI in healthcareinnovation adoption and diffusion
Guodong Gordon Gao
Guodong Gordon Gao
Johns Hopkins University
AI in HealthcareQuality TransparencymHealthBig Data