Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study evaluates the efficacy and bias of large language models (LLMs) in clinical research retrieval. Leveraging Cochrane Reviews as a gold standard, we systematically examined the impact of model type, user persona prompting, and sample size on recall through multi-model comparisons and regression analysis. Results indicate that ChatGPT achieves optimal recall performance, with researcher personas significantly outperforming patient personas. Notably, sample size emerged as the sole independent significant predictor of retrieval outcomes, revealing an inherent LLM bias toward large-sample trials. By quantifying these critical determinants, this work provides empirical evidence to inform the optimization of intelligent retrieval systems for evidence-based medicine, highlighting the necessity of addressing sample-size-related biases in LLM-assisted clinical search tasks.
📝 Abstract
Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studies and the factors driving their selection. In this study, we evaluated three general-purpose LLM chatbots: Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5. We prompted the models with clinical questions adapted from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews, simulating patient, clinician, and evidence-synthesis researcher roles. Each chatbot was queried under each user role with four independent repetitions, yielding 720 responses. Each chatbot was asked to support its answers with primary clinical citations, which we benchmarked against the included and excluded study sets of the Cochrane reviews. On average, a chatbot response retrieved 39.2% $\pm$ 29.8% of Cochrane included studies, while citing 5.0% $\pm$ 9.4% of excluded studies. Recall of Cochrane included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%; $p=2.0\times10^{-5}$). The researcher role yielded higher recall than the clinician or patient roles (42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%; $p=2.0\times10^{-5}$). Controlling for publication year, citations per year, and open-access status, sample size was the only independently significant predictor of retrieval (odds ratio 1.80 per 1-unit increase in log sample size, 95% CI 1.37-2.36, $p=2.34\times10^{-5}$). These findings suggest that while LLM chatbots can retrieve some studies identified by expert reviewers, their performance varies by model and user role, and they exhibit a bias toward clinical trials with larger sample sizes.
Problem

Research questions and friction points this paper is trying to address.

LLM chatbots
study retrieval
clinical questions
retrieval bias
evidence quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Study Retrieval Evaluation
Sample Size Bias
User Role Simulation
Cochrane Benchmarking
Evidence Synthesis
Q
Qingfang Liu
National Institute on Drug Abuse Intramural Research Program, National Institutes of Health, Baltimore, MD, USA
Q
Qiao Jin
National Library of Medicine, National Institutes of Health, Bethesda, MD, USA
J
Joe D. Menke
School of Information Sciences, University of Illinois Urbana-Champaign, Champaign, IL, USA
T
Thorsten Kahnt
National Institute on Drug Abuse Intramural Research Program, National Institutes of Health, Baltimore, MD, USA
Zhiyong Lu
Zhiyong Lu
Senior Investigator, NLM; Adjunct Professor of CS, UIUC
BioNLPBiomedical InformaticsMedical AIArtificial Intelligence