Trustworthy Protein-Ligand Binding Affinity Prediction via Reliability-Aware Multi-Engine Fusion

๐Ÿ“… 2026-07-20
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the lack of reliable confidence estimation in existing proteinโ€“ligand binding affinity prediction methods, which hinders the assessment of result credibility across different docking engines. The authors propose RELIABLE-BA, a novel framework that integrates evidential theory with context-aware reliability modeling. By employing a Normal-Inverse-Gamma distribution to construct an evidential expert model, the method dynamically evaluates the reliability of each docking engine within specific molecular contexts and introduces a closed-form uncertainty fusion algorithm enabling interpretable decomposition of predictive uncertainty. Evaluated on PDBBind and BDB2020+ benchmarks, RELIABLE-BA achieves both high-accuracy point predictions and well-calibrated uncertainty estimates. Validation on SARS-CoV-2 Mpro and 5HT2A targets demonstrates up to a 25% reduction in prediction error for high-confidence subsets, offering a robust pathway for AI-assisted drug discovery.
๐Ÿ“ Abstract
Accurate protein-ligand binding affinity prediction is central to computational drug discovery, yet modern docking engines frequently disagree without indicating which prediction to trust. Consensus scoring and ensemble methods improve mean accuracy but treat all predictions identically without interpretable confidence measures or uncertainty decomposition, ignoring the chemical context of each protein-ligand pair. To address this limitation, we introduce RELIABLE-BA (RELIABiLity-aware Evidential fusion for Binding Affinity), an evidential framework for multi-engine binding affinity prediction. Our model comprises three steps: (1) modeling each engine as an evidential expert via Normal-Inverse-Gamma distributions, (2) scaling epistemic uncertainty through learned reliability from molecular context while preserving each expert's predictive mean, and (3) fusing experts through closed-form aggregation that captures both individual uncertainty and inter-engine disagreement. Experiments on the PDBBind and BDB2020+ benchmarks demonstrate competitive point prediction with substantially improved uncertainty calibration, and additional validation on the SARS-CoV-2 Mpro dataset and 5HT2A receptor demonstrates applicability to clinically relevant drug targets. Crucially, these uncertainty estimates enable reliable filtering of protein-ligand pairs, reducing prediction error by up to 25% when retaining only high-confidence pairs. To our knowledge, RELIABLE-BA is the first multi-engine binding affinity prediction framework to combine evidential fusion with context-dependent reliability, offering a principled path toward trustworthy AI-guided drug discovery. Our code is publicly available at https://github.com/yongchand/RELIABLE-BA.
Problem

Research questions and friction points this paper is trying to address.

protein-ligand binding affinity
uncertainty calibration
docking engines
trustworthy prediction
computational drug discovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

evidential fusion
reliability-aware learning
binding affinity prediction
uncertainty calibration
multi-engine docking
๐Ÿ’ผ Related Jobs
No related jobs found.
Y
Yongchan Hong
Department of Quantitative & Computational Biology, University of Southern California, Los Angeles, California, USA
Defu Cao
Defu Cao
Peking University; MBZUAI; University of Southern California; Caltech
Time SeriesFoundation ModelMachine LearningCausal InferenceLLM
W
Wenjin Liu
Department of Quantitative & Computational Biology, University of Southern California, Los Angeles, California, USA
T
Thomas Ku
Department of Quantitative & Computational Biology, University of Southern California, Los Angeles, California, USA
J
Jordy Homing Lam
Department of Chemistry, University of California, Berkeley, Berkeley, California, USA
E
Emily Nguyen
Department of Computer Science, University of Southern California, Los Angeles, California, USA
Willie Neiswanger
Willie Neiswanger
Assistant Professor of Computer Science, University of Southern California
Machine LearningStatisticsOptimizationSequential Decision MakingAI-for-Science
Vsevolod Katritch
Vsevolod Katritch
University of Southern California
GPCRcomputational structural biologycomputer aided drug designstructural bioinformatics
Yan Liu
Yan Liu
Professor, Director of Machine Learning Center@USC
ML for Time SeriesExplainable AIPhysics-informed AIAI for health careAI for Sustainability