Learning from Literature: Integrating LLMs and Bayesian Hierarchical Modeling for Oncology Trial Design

📅 2026-02-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of modern oncology trials, which often suffer from hypothesis bias and underpowered sample sizes due to reliance on incomplete literature abstracts, leading to false-positive or false-negative conclusions. To overcome this, we propose the LEAD-ONC framework, which uniquely integrates large language models (LLMs) with Bayesian hierarchical modeling to automatically extract baseline characteristics from unstructured clinical trial reports, reconstruct individual patient data, and generate survival prediction distributions for target populations. Applied to five phase III trials of first-line treatment in non-small cell lung cancer, our approach identified three clinically interpretable subgroups and predicted a 2.8-month difference in median overall survival (95% credible interval: –2.0 to 7.6) between immunotherapy monotherapy and combination regimens in a mixed-histology population, with a 45% probability of achieving more than three months of benefit—substantially enhancing the precision and prospectiveness of trial design.

Technology Category

Application Category

📝 Abstract
Designing modern oncology trials requires synthesizing evidence from prior studies to inform hypothesis generation and sample size determination. Trial designs based on incomplete or imprecise summaries can lead to misspecified hypotheses and underpowered studies, resulting in false positive or negative conclusions. To address this challenge, we developed LEAD-ONC (Literature to Evidence for Analytics and Design in Oncology), an AI-assisted framework that transforms published clinical trial reports into quantitative, design-relevant evidence. Given expert-curated trial publications that meet prespecified eligibility criteria, LEAD-ONC uses large language models to extract baseline characteristics and reconstruct individual patient data from Kaplan-Meier curves, followed by Bayesian hierarchical modeling to generate predictive survival distributions for a prespecified target trial population. We demonstrate the framework using five phase III trials in first-line non-small-cell lung cancer evaluating PD-1 or PD-L1 inhibitors with or without CTLA-4 blockade. Clustering based on baseline characteristics identified three clinically interpretable populations defined by histology. For a prospective randomized trial in the mixed-histology population comparing mono versus dual immune checkpoint inhibition, LEAD-ONC projected a modest median overall survival difference of 2.8 months (95 percent credible interval -2.0 to 7.6) and an estimated probability of at least a 3-month benefit of approximately 0.45. As LEAD-ONC remains under active development, these results are intended as preliminary demonstrations of the frameworks potential to support evidence-driven oncology trial design rather than definitive clinical conclusions.
Problem

Research questions and friction points this paper is trying to address.

oncology trial design
evidence synthesis
hypothesis generation
sample size determination
clinical trial literature
Innovation

Methods, ideas, or system contributions that make the work stand out.

large language models
Bayesian hierarchical modeling
individual patient data reconstruction
Kaplan-Meier curve digitization
oncology trial design
🔎 Similar Papers
No similar papers found.
G
Guannan Gong
Yale Cancer Center, Yale School of Medicine, New Haven, CT, USA
S
Satrajit Roychoudhury
Pfizer Inc, New York, NY, USA
A
Allison Meisner
Public Health Sciences Division, Fred Hutchinson Cancer Center
L
Lajos Pusztai
Yale Cancer Center, Yale School of Medicine, New Haven, CT, USA
S
Sarah B Goldberg
Yale Cancer Center, Yale School of Medicine, New Haven, CT, USA
W
Wei Wei
Yale Cancer Center, Yale School of Medicine, New Haven, CT, USA