Multi-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC)

📅 2026-07-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing rule-based approaches struggle to comprehensively capture the multidimensional nature of disease severity in electronic health records. This work proposes MOSAIC, a novel framework that, for the first time, leverages a multi-agent large language model system to assess phenotypic severity in type 2 diabetes by integrating key dimensions such as biomarkers, β-cell function, and social determinants of health, thereby overcoming the rigidity of fixed-rule systems. Trained and validated against established clinical scoring criteria (DCSI, DiSSCo, Cooper) and real-world clinical outcomes, the open-source implementation demonstrates strong agreement with its closed-source counterpart (weighted kappa = 0.773). Stratified severity levels significantly differentiate all-cause mortality and complication risk (log-rank p < 0.001), outperforming rule-based baselines.
📝 Abstract
Background: Disease severity is a multidimensional construct difficult to capture with rule-based approaches in Electronic Healthcare Records (EHR). Agentic large language model (LLM) systems could synthesise clinical evidence and reason over EHRs, but remain unevaluated for this task. Methods: MOSAIC is a two-phase agentic LLM framework for severity phenotyping, using type 2 diabetes (T2D) as a proof-of-concept. MOSAIC was evaluated on a synthetic cohort (SyntheticMass; open-weight N = 4,886; closed-weight N = 200) against three algorithmic ground truths (DCSI, DiSSCo, Cooper) and against all-cause mortality and incident complications. Open-weight (locally deployable) and proprietary pipelines were also compared. Results: The generated framework spanned domains absent from the comparators, including biomarker-based glycaemic staging, beta-cell function, and social determinants of health. Open-weight MOSAIC matched the proprietary pipeline (closed- vs open-weight weighted kappa = 0.773) and reached moderate agreement with Cooper (kappa = 0.597) and DCSI (kappa = 0.534) and fair agreement with DiSSCo (kappa = 0.320). Agent-based (Type 1) tiers showed significant separation of all-cause mortality (log-rank p < 0.001; crude hazard ratios 1.6-2.4 for non-Baseline tiers), with non-monotonic separation at the upper tiers, and an inverse gradient for incident complications (log-rank p < 0.001) consistent with depletion of susceptibles. Agentic classification also diverged from deterministic execution of the same rubric (MOSAIC Frozen; kappa = 0.428), indicating reasoning beyond fixed rules. Conclusion: MOSAIC shows agentic LLM systems can generate and apply clinically meaningful severity phenotypes from structured EHR data in T2D. Extending it to other diseases with similarly multidimensional severity warrants further research.
Problem

Research questions and friction points this paper is trying to address.

disease severity
Electronic Healthcare Records
severity phenotyping
type 2 diabetes
multidimensional assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

agentic LLM
severity phenotyping
electronic health records
multidimensional assessment
open-weight models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Manuela Del Castillo Suero
Department of Drug Design and Pharmacology, University of Copenhagen, Copenhagen, Denmark
A
Arnault-Quentin Vermillet
Department of Drug Design and Pharmacology, University of Copenhagen, Copenhagen, Denmark
N
Nicole Sonne Heckmann
Department of Drug Design and Pharmacology, University of Copenhagen, Copenhagen, Denmark
D
Darmendra Ramcharran
GSK, Providence, RI, USA; School of Pharmacy, University of Rhode Island, Kingston, RI, USA
M
Maurizio Sessa
Department of Drug Design and Pharmacology, University of Copenhagen, Copenhagen, Denmark