Medical Causal Hypothesis Verification with Large Language Models

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究评估了大型语言模型在验证医学因果假设时的准确性,通过提出一个评价框架并使用17个医学因果假设测试8个模型,揭示了这些模型在提供有效科学证据上的局限性。
📝 Abstract
The growing use of large language models (LLMs) for search and information retrieval underscores the need to evaluate their reliability in high-stakes domains such as healthcare. Although LLMs can effectively answer questions about diseases, symptoms, and treatments, their ability to accurately assess causal relationships and ground their conclusions in verified scientific evidence remains unclear. Here, we present a preliminary, small-scale study that investigates the accuracy of LLMs in evaluating causal medical claims and supporting them with peer-reviewed research. We propose an evaluation framework for causal hypothesis verification that can be used to systematically track the performance of existing and future LLMs. We assess the performance of eight LLMs on 17 medical causal hypotheses to evaluate whether they can reliably verify these hypotheses using scientific evidence from the literature. We systematically annotate the scientific evidence they provide according to six criteria (a total of 1,067 annotation points) and assess them with nine evaluation metrics. Our analysis shows that while LLMs exhibit strong recall, they often perform poorly at providing valid scientific articles and evidence for support and at rejecting unsupported hypotheses. These findings highlight a critical limitation of current LLMs, as they cannot yet be trusted fully to verify causal relationships from the biomedical literature. This work underscores the need for rigorous evaluation before using LLMs for search and retrieval in healthcare settings.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Causal Hypothesis Verification
Healthcare
Scientific Evidence
Peer-Reviewed Research
Innovation

Methods, ideas, or system contributions that make the work stand out.

causal hypothesis verification
evaluation framework
large language models (LLMs)
medical claims
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
S
Safiyyah Ahmed
University of Illinois Chicago
A
Abrar Ansari
University of Illinois Chicago
M
Md Aminul Islam
University of Illinois Chicago
Elena Zheleva
Elena Zheleva
Associate Professor of Computer Science, University of Illinois Chicago
Machine learningCausal inferenceData scienceGraph miningPrivacy