When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文研究了医学AI评估中评分标准未能捕捉临床相关错误的问题,通过开发错误注入流程和分类方法,揭示了现有评估方法的盲点。
📝 Abstract
Hallucinations can undermine clinician trust in LLMs, making it important that evaluation methods capture clinically relevant errors. Rubric-based evaluation has become the leading approach for assessing LLMs in medicine, but it is unclear whether rubric scores reflect such errors. We first study this in a controlled setting using MedHallu, finding that more specific rubrics better distinguish correct from hallucinated responses. To test this systematically, we develop a taxonomy of medical hallucination types and a clinician-validated error-injection pipeline that creates matched correct and error-injected responses. Across HealthBench, HealthBench Professional, and LiveMedBench, our clinically relevant hallucinations are missed by rubrics, often leaving scores unchanged. We find that rubrics are most effective when explicitly checking facts, and are less effective for additional or unexpected errors they do not anticipate. A preliminary retrieval-based factuality check recovers some of the rubric-blind errors, suggesting a complementary approach. These findings reveal systematic blind spots in current medical evaluation of LLMs and suggest that rubric scores alone are insufficient to establish clinical reliability, potentially undermining clinician trust and confidence in clinical deployment.
Problem

Research questions and friction points this paper is trying to address.

Hallucinations
Medical AI Evaluation
Rubric-based evaluation
Clinician trust
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hallucinations
Rubric-based Evaluation
Medical AI
Error-injection Pipeline
Factuality Check
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Griffin Farrow
University of Oxford
L
Lily Sijia Li
University of Oxford
J
Jack Johnson
University of Oxford
Tingyan Wang
Tingyan Wang
University of Oxford
Philip Torr
Philip Torr
Professor, University of Oxford
Department of Engineering
W
William Bolton
University of Oxford
F
Fabio J. Fehr
University of Oxford