MamaBench: Benchmarking LLM Robustness in Maternal and Child Health Diagnosis through Counterfactual Clinical Perturbation

📅 2026-07-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current medical benchmarks inadequately assess large language models’ ability to distinguish between clinically similar cases that require different interventions, often leading to an overestimation of their robustness. To address this gap, this work introduces MamaBench, the first counterfactual clinical diagnostic benchmark focused on maternal and child health, alongside an Evidence-Anchored Retrieval-Augmented Generation (EA-RAG) framework. EA-RAG employs a three-stage, evidence-coverage-driven retrieval mechanism—comprising clinical parameter extraction, coverage auditing, and contrastive sub-query generation—that overcomes the limitations of conventional similarity-based aggregation. Evaluated on Claude Sonnet 4.6, the approach achieves a robust accuracy of 65.0% and reduces the Bias Trap Rate (BTR) to 20.3%, a 5.5-percentage-point improvement over the baseline without compromising base performance, revealing that current models still fail in approximately 20% of counterfactual clinical reasoning scenarios.
📝 Abstract
Large language models achieve strong scores on medical benchmarks, yet these benchmarks evaluate each question in isolation, providing no measure of whether a system can distinguish clinically similar presentations requiring different interventions. We introduce MamaBench, the first counterfactual benchmark for maternal and paediatric AI: 434 expert-authored clinical narratives in 217 pairs across 371 pathologies, evaluated via the Bias Trap Rate (BTR), the conditional probability that a model fails the counterfactual given success on the base case. We propose Evidence-Anchored RAG (EA-RAG), a three-stage retrieval method that replaces aggregate similarity with an evidence coverage objective through clinical parameter extraction, coverage auditing, and contrastive sub-queries. Across eight configurations of four frontier LLMs, base accuracy overstates robust accuracy by 16-28 percentage points in every model. EA-RAG achieves 20.3% BTR and 65.0% robust accuracy on Claude Sonnet 4.6, a 5.5 percentage point BTR reduction without degrading base accuracy. The residual 20% BTR confirms that counterfactual robustness in clinical AI remains an open challenge. Keywords: counterfactual evaluation, clinical AI, maternal healthcare, retrieval-augmented generation, diagnostic robustness
Problem

Research questions and friction points this paper is trying to address.

counterfactual evaluation
clinical AI
diagnostic robustness
maternal healthcare
Innovation

Methods, ideas, or system contributions that make the work stand out.

counterfactual evaluation
Evidence-Anchored RAG
Bias Trap Rate
clinical AI
diagnostic robustness
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Thanni Adewuyi
Helpmum Africa, University of Ibadan
A
Anuoluwa Sotome
Helpmum Africa
S
Samuel Okoko
Helpmum Africa
A
Angel Ezendu
Helpmum Africa
O
Oluwafunke Akinbuwa
Helpmum Africa
O
Oluwaseun Odunsi
Helpmum Africa
O
Oluwasegun Oguntuase
Helpmum Africa
O
Oluwadarasimi Oguntuase
Helpmum Africa
I
Ifeoma Nwabueze
Helpmum Africa
A
Abiodun Adereni
Helpmum Africa