๐ค AI Summary
Existing protein datasets are largely confined to static structures or monomeric dynamic trajectories, making it difficult to model the formation process of multichain protein complexes. To address this gap, this work presents DynaPPI, the first systematically constructed dataset comprising full-timescale molecular dynamics trajectories that capture the transition from unbound monomers to bound complexes, thereby bridging the divide between static structures and dynamic interactions. Leveraging this dataset, we introduce a diffusion-based model capable of learning the dynamic assembly pathways of protein complexes and achieving high-accuracy prediction of three-dimensional structures for previously unseen proteinโprotein complexes. This study establishes a foundational data resource and methodological framework for AI-driven interactome research.
๐ Abstract
Diffusion models have been widely explored in protein backbone generation due to their powerful generation capabilities.However, in today's AI-driven biological research, predicting the structure of unknown multi-chain protein aggregates (called "complexes" in biology) remains an unsolved challenge.This is because existing static or dynamic protein datasets focus solely on static snapshots or single-entity trajectories, neglecting the dynamic process of multiple monomers forming complexes.To alleviate this dilemma, we present DynaPPI, a dynamic protein dataset comprising molecular dynamics (MD) trajectories of protein complex formation from dissociated chains to the bound state, as a pivotal resource to bridge the gap between static structural biology and the inherently temporal nature of dynamic molecular interactions.Benefiting from this dataset, diffusion models can explicitly learn the dynamic binding trajectories of known complexes and accurately predict the structures of unknown complexes based on their diverse generative properties, thereby further catalyzing AI-driven structural biology and protein interactomics.