binding affinity prediction

Builds predictive models that estimate binding affinity between ligands and targets using structural, biophysical, and machine-learning features.

bindingaffinityprediction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.13
Aug 01, 2026Aug 01, 2026
Career
Value
No comparison yet
$200K/year
Aug 01, 2026Aug 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Binding Affinity Prediction: From Conventional to Machine Learning-Based Approaches

Sep 30, 2024
XL
Xuefeng Liu
🏛️ University of Chicago | Argonne National Laboratory | Data Science Institute

Protein–ligand binding affinity prediction remains hindered by limited generalizability, poor interpretability, and insufficient robustness under low-data regimes. Method: This study conducts a systematic review and empirical evaluation of affinity prediction approaches—including physics-based models, traditional machine learning (e.g., RF, XGBoost), and deep learning architectures (e.g., GCN, SE(3)-Transformer, GNNs)—integrating molecular docking, 3D conformation generation, sequence/structure embeddings, and energy-based features. Contribution/Results: We present the first comprehensive analysis of methodological evolution, identify three critical bottlenecks, and propose a novel paradigm combining multi-scale representation fusion with physics-informed learning. We establish data quality, negative sample construction, and cross-target transfer as key determinants of performance. Our framework achieves R² = 0.82 on PDBbind v2020 and delivers a principled, AI-driven methodology for drug discovery.

Advancing AI-driven in silico models for therapeuticsComparing conventional and machine learning approachesPredicting protein-ligand binding affinity strength

This work proposes a coarse-grained, diffusion-free multimodal framework for drug–protein binding affinity prediction, addressing the high computational cost of existing all-atom diffusion models that hinders large-scale screening. By representing binding pockets using only protein Cβ atoms and ligand heavy atoms, the method integrates COATI-3 molecular encodings with ESM-2 protein embeddings and incorporates diffusion-free conformational optimization, affinity likelihood prediction, calibratable uncertainty estimation, and continual learning. The approach achieves conformation generation accuracy comparable to diffusion-based models while improving Pearson correlation coefficients for affinity prediction by approximately 20%. Furthermore, it demonstrates a six-fold increase in molecular optimization efficiency over greedy strategies, offering a compelling balance between computational efficiency and predictive reliability.

all-atom diffusionbinding affinity predictioncomputational bottleneck

This study addresses the limited fine-grained evaluation of binding sites and non-covalent interactions in existing protein–ligand binding models, which primarily focus on binary binding prediction and affinity estimation. To bridge this gap, the authors introduce InteractBind, a dataset comprising approximately 100,000 protein–ligand complexes, and propose a novel residue-to-atom-level binding site localization task based on six distinct types of non-covalent interactions. They further devise a data splitting strategy that accounts for both binding affinity and protein similarity to rigorously assess model generalization and physical interpretability. Benchmarking eight state-of-the-art models under a unified protocol reveals strong performance in binding prediction but consistently limited accuracy in binding site localization, with substantial performance variation across different interaction types.

binding site localizationmodel interpretabilitymolecular recognition

Learning to Align Molecules and Proteins: A Geometry-Aware Approach to Binding Affinity

Sep 24, 2025
MR
Mohammadsaleh Refahi
🏛️ Drexel University

Accurate and generalizable drug–target binding affinity prediction remains challenging, as existing deep learning models often rely on simplistic feature concatenation without geometric constraints, limiting extrapolation beyond training chemical spaces. To address this, we propose a geometry-aware molecular alignment framework that jointly models ligand–protein conditional dependencies via Feature-wise Linear Modulation (FiLM) and enforces metric consistency in the embedding space through an RBF-based regression head coupled with a triplet distance loss. This lightweight architecture explicitly incorporates structural priors into representation learning. Evaluated on the Therapeutics Data Commons DTI-DG benchmark, our method achieves state-of-the-art performance. Ablation studies confirm the efficacy of each component, while cross-domain experiments demonstrate substantial improvements in out-of-distribution generalization and model interpretability.

Improving generalization across chemical space and timeIncorporating geometric regularization for robust affinity predictionPredicting drug-target binding affinity to accelerate drug discovery

Deep Learning for Protein-Ligand Docking: Are We There Yet?

May 23, 2024
AM
Alex Morehead
🏛️ University of Missouri

This work addresses the generalization bottleneck of deep learning (DL) methods for protein–ligand docking in realistic scenarios, focusing on three key challenges: (1) pocket-agnostic docking (i.e., without prior binding-site annotation), (2) multi-ligand cooperative docking (e.g., cofactor binding), and (3) docking into predicted apo-protein structures (critical for novel targets). To rigorously evaluate cross-domain generalization under these conditions, we introduce PoseBench—the first comprehensive, application-oriented benchmark for real-world docking—and publicly release it with support for both single- and multi-ligand evaluation. Methodologically, our approach integrates deep structural modeling, physics-informed loss functions, complex-aware clustering during training, and generative structural refinement. Experiments demonstrate that DL-based methods consistently outperform traditional algorithms overall; however, most exhibit limited generalization to multi-ligand settings. Crucially, incorporating physics-guided constraints significantly enhances robustness—particularly for apo-protein docking and unknown-pocket scenarios.

Assessing docking with predicted protein structures.Benchmarking methods for multi-ligand binding scenarios.Evaluating deep learning for protein-ligand docking.

Latest Papers

What's happening recently
View more

Accurate prediction of protein–ligand binding affinity is hindered by data scarcity, experimental heterogeneity, and conformational dependence. This work proposes a novel paradigm that leverages a multi-head frozen stochastic atom graph encoder to generate diverse structural representations, integrates explicit physicochemical interaction fingerprints, and employs a heterogeneous regressor combining neural networks and tree-based models. To enhance generalization and robustness, the approach avoids end-to-end training and instead adopts a validation-based non-negative fusion strategy. Evaluated on the GEMS-reconstructed PDBbind 2020R1 similarity-isolated split and the CASF-2016 benchmark, the method demonstrates consistently strong and stable predictive performance.

binding affinity predictioncomputational chemistrymolecular modeling

This study systematically evaluates the reliability of the AI model Boltz-2 in predicting protein–ligand complex structures and binding affinities for drug discovery. Leveraging a large-scale dataset of real-world drug targets, it presents the first independent validation of Boltz-2’s “co-folding” strategy, benchmarking its performance against conventional molecular docking and physics-based ESMACS free energy calculations. The results indicate that while Boltz-2 shows promise for rapid initial screening, it exhibits significant limitations in structural convergence and binding free energy accuracy, rendering it insufficient for lead compound identification. This work underscores the critical need to integrate physics-based methods in validating AI-driven approaches and establishes an essential empirical benchmark for future AI-enabled drug design efforts.

AI reliabilitybinding affinityBoltz-2

This work addresses a central challenge in drug discovery: designing small-molecule ligands that bind target protein pockets with high affinity and specificity. The authors propose DBMol, a novel framework that, for the first time, leverages structure prediction models such as AlphaFold-3 and Boltz-2 as unsupervised optimization signals. By integrating gradient-based affinity optimization with a flow-matching generative model, DBMol enables iterative refinement from an initial molecule while ensuring chemically valid structures. Notably, the method operates without requiring reference ligands and generates molecules exhibiting high predicted binding affinity, target specificity, and structural diversity. Experimental results demonstrate that DBMol significantly outperforms existing unconditional generative approaches in both protein pocket coverage and molecular generation quality.

binding affinityde novo drug designprotein-ligand binding

This work addresses the limitation of existing protein–ligand binding affinity prediction methods that rely predominantly on static structures and overlook molecular flexibility and binding-induced conformational changes. The authors propose a Curvature-aware Potential Energy Surface (CPES) graph neural network, which introduces—for the first time—the local principal curvature spectrum derived from the Hessian matrix of the potential energy surface as a dynamic descriptor within a deep learning framework. By integrating this dynamic information with static geometric features, the model employs a spectral cross-attention mechanism to capture conformational differences before and after binding. It further combines geometry-aware message passing with soft clustering to learn multi-scale interactions and uses bidirectional cross-attention to fuse dynamic and static representations for affinity regression. The method achieves significantly improved prediction performance across multiple benchmarks and offers physically interpretable insights into conformational changes.

conformational changesgeometric deep learningmolecular flexibility

Traditional protein–ligand affinity prediction models are limited by their inability to accurately represent binding contact maps, which compromises virtual screening efficiency. This work proposes PLANET v2.0, a novel framework that integrates graph neural networks with mixture density networks through a multi-objective training strategy. It innovatively employs two Gaussian mixture models to probabilistically capture the spatial distribution and energetic characteristics of non-covalent interactions, enabling joint prediction of binding poses and affinities in the form of probability densities. The final affinity score is derived via mathematical expectation. On the CASF-2016 benchmark, PLANET v2.0 significantly outperforms both its predecessor PLANET and Glide SP across all four tasks—scoring, ranking, docking, and screening—and demonstrates exceptional robustness on a large-scale commercial dataset.

Binding Mode PredictionContact Map RepresentationProtein-Ligand Affinity Prediction

Hot Scholars

JG

Jiaqi Guan

Research Scientist, ByteDance AML AI4Science
Machine LearningGenerative ModelsComputational ChemistryComputational Biology
XZ

Xiangxiang Zeng

Deparment of Computer Science, Hunan University
Computational IntelligenceAI4ScienceAI for Drug Discovery
JM

Jianzhu Ma

Tsinghua University
Machine LearningComputational BiologyBioinformatics
TF

Tianfan Fu

Nanjing University
AI for DrugAI for ScienceLarge Language Model