Institution profile

HelpMum

Research institution
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

MamaBench: Benchmarking LLM Robustness in Maternal and Child Health Diagnosis through Counterfactual Clinical Perturbation

Jul 15, 2026

Current medical benchmarks inadequately assess large language models’ ability to distinguish between clinically similar cases that require different interventions, often leading to an overestimation of their robustness. To address this gap, this work introduces MamaBench, the first counterfactual clinical diagnostic benchmark focused on maternal and child health, alongside an Evidence-Anchored Retrieval-Augmented Generation (EA-RAG) framework. EA-RAG employs a three-stage, evidence-coverage-driven retrieval mechanism—comprising clinical parameter extraction, coverage auditing, and contrastive sub-query generation—that overcomes the limitations of conventional similarity-based aggregation. Evaluated on Claude Sonnet 4.6, the approach achieves a robust accuracy of 65.0% and reduces the Bias Trap Rate (BTR) to 20.3%, a 5.5-percentage-point improvement over the baseline without compromising base performance, revealing that current models still fail in approximately 20% of counterfactual clinical reasoning scenarios.

0 citationsRead paper
Recent publications

Latest Papers

MamaBench: Benchmarking LLM Robustness in Maternal and Child Health Diagnosis through Counterfactual Clinical Perturbation

Jul 15, 2026

Current medical benchmarks inadequately assess large language models’ ability to distinguish between clinically similar cases that require different interventions, often leading to an overestimation of their robustness. To address this gap, this work introduces MamaBench, the first counterfactual clinical diagnostic benchmark focused on maternal and child health, alongside an Evidence-Anchored Retrieval-Augmented Generation (EA-RAG) framework. EA-RAG employs a three-stage, evidence-coverage-driven retrieval mechanism—comprising clinical parameter extraction, coverage auditing, and contrastive sub-query generation—that overcomes the limitations of conventional similarity-based aggregation. Evaluated on Claude Sonnet 4.6, the approach achieves a robust accuracy of 65.0% and reduces the Bias Trap Rate (BTR) to 20.3%, a 5.5-percentage-point improvement over the baseline without compromising base performance, revealing that current models still fail in approximately 20% of counterfactual clinical reasoning scenarios.

0 citationsRead paper