Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation

📅 2025-05-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliability bottleneck in large language models (LLMs) for clinical note generation (CNG), stemming from response variability. We propose the first multidimensional reliability evaluation framework tailored to healthcare settings, assessing string-level consistency, semantic consistency, and semantic correctness. We systematically evaluate 12 open- and closed-source LLMs using repeated sampling, prompt-consistency testing, and automated metrics (BLEU, BERTScore, natural language inference). All evaluations are validated by clinical domain experts. Results show that all models achieve >92% semantic consistency; Llama-70B attains the highest overall reliability—significantly outperforming most commercial LLMs. Moreover, smaller-parameter open-source models demonstrate superior accuracy, robustness, and compliance with local deployment and privacy requirements. This work establishes a reproducible, interpretable, and privacy-preserving evaluation paradigm for clinical deployment of LLM-driven CNG systems.

Technology Category

Application Category

📝 Abstract
Due to the legal and ethical responsibilities of healthcare providers (HCPs) for accurate documentation and protection of patient data privacy, the natural variability in the responses of large language models (LLMs) presents challenges for incorporating clinical note generation (CNG) systems, driven by LLMs, into real-world clinical processes. The complexity is further amplified by the detailed nature of texts in CNG. To enhance the confidence of HCPs in tools powered by LLMs, this study evaluates the reliability of 12 open-weight and proprietary LLMs from Anthropic, Meta, Mistral, and OpenAI in CNG in terms of their ability to generate notes that are string equivalent (consistency rate), have the same meaning (semantic consistency) and are correct (semantic similarity), across several iterations using the same prompt. The results show that (1) LLMs from all model families are stable, such that their responses are semantically consistent despite being written in various ways, and (2) most of the LLMs generated notes close to the corresponding notes made by experts. Overall, Meta's Llama 70B was the most reliable, followed by Mistral's Small model. With these findings, we recommend the local deployment of these relatively smaller open-weight models for CNG to ensure compliance with data privacy regulations, as well as to improve the efficiency of HCPs in clinical documentation.
Problem

Research questions and friction points this paper is trying to address.

Evaluating LLM reliability in clinical note generation
Assessing consistency and accuracy of LLM-generated notes
Comparing performance of open-weight and proprietary LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluates 12 LLMs for clinical note reliability
Measures consistency, semantic accuracy, correctness
Recommends local deployment of open-weight models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kristine Ann Carandang
Analytics, Computing & Complex Systems Laboratory, Asian Institute of Management
J
Jasper Meynard P. Arana
Analytics, Computing & Complex Systems Laboratory, Asian Institute of Management
E
Ethan Robert A. Casin
Analytics, Computing & Complex Systems Laboratory, Asian Institute of Management
C
Christopher P. Monterola
Analytics, Computing & Complex Systems Laboratory, Asian Institute of Management
D
Daniel Stanley Y. Tan
Analytics, Computing & Complex Systems Laboratory, Asian Institute of Management
J
Jesus Felix B. Valenzuela
Analytics, Computing & Complex Systems Laboratory, Asian Institute of Management
C
Christian M. Alis
Analytics, Computing & Complex Systems Laboratory, Asian Institute of Management