🤖 AI Summary
This study addresses the suboptimal performance of large language models (LLMs) in causality assessment for automated pharmacovigilance and the absence of effective methods for optimizing inference hyperparameters such as temperature. To this end, the authors propose a Gaussian process–based Bayesian optimization framework that systematically tunes temperature for LLM-based causal inference, incorporating a novel weighted consistency metric—particularly the Entropy-Weighted Agreement Consistency Score (EWACS). Evaluated on individual case safety reports from FAERS using GPT-5.2, chain-of-thought prompting, and four consistency measures, the approach significantly improves agreement between model predictions and expert judgments from 45.0% to 72.0%, with a 42.9-percentage-point gain in the “suspected” category. These results demonstrate that optimal temperature is highly context-dependent, precluding a universal setting, and substantially enhance the practical utility of LLMs in regulatory pharmacovigilance applications.
📝 Abstract
Background: Growing individual case safety report (ICSR) volumes have intensified demand for scalable automated causality assessment. Large Language Models (LLMs) show promise, yet performance on clinically demanding tasks remains suboptimal and inference-time hyperparameter optimization has not been investigated. Objective: To develop a Gaussian Process (GP)-compatible optimization objective and investigate whether temperature optimization improves LLM-expert agreement on Naranjo causality assessment of FAERS ICSRs. Methods: Expert causality assessments were performed on 723 stratified FAERS cases. OpenAI's GPT-5.2 was evaluated using chain-of-thought (CoT) prompting. Four composite metrics were developed: Weighted Cosine Similarity (WCS), Information-Weighted Agreement Score (IWAS), Entropy-Weighted Agreement and Cosine Similarity Score (EWACS), and Consensus-Weighted Cosine Similarity (CWCS) and Bayesian optimization using a GP surrogate with Probability of Improvement (PoI) acquisition was applied across temperature [0, 2]. Results: GPT-5.2 outperformed prior biomedical LLMs at baseline (T = 0), achieving 74.1% agreement on question 5 and 65.4% on question 10 of Naranjo algorithm. Entropy analysis identified these as the sole informative optimization targets. Temperature showed no systematic population-level effect (\b{eta} = 0.002, p = 0.959). EWACS-guided Bayesian optimization improved causality classification agreement from 45.0% to 72.0% (+27 pp), with the largest gain in Doubtful cases (+42.9 pp). Conclusion: EWACS was identified as the optimal GP-compatible metric. The absence of a universal temperature optimum indicates LLM performance is driven primarily by ICSR content, yet case-specific temperature selection produced meaningful improvements, supporting temperature optimization for LLM-assisted pharmacovigilance.