FoldKit: A Python library for efficient storage and retrieval of co-folding predictions

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决AlphaFold 3共折叠实验产生的大量数据存储问题,开发了Python库FoldKit,通过高效压缩和结构化处理这些数据,并提供便捷的程序访问接口。
📝 Abstract
AlphaFold 3 (AF3) enables structure prediction of biomolecular complexes through co-folding multiple interacting molecules, making it increasingly useful for de novo protein design and for large-scale studies of protein-protein, protein-peptide, and other biomolecular interactions. However, systematic co-folding experiments can produce large volumes of output data, particularly when multiple random seeds and samples are generated for each input complex. We introduce FoldKit, a Python package for efficient storage and analysis of large-scale AF3 co-folding results. FoldKit converts raw AF3 outputs into a compact, structured representation while preserving the metadata needed for downstream analysis. The FoldKit Python library provides convenient programmatic access to global, single chain, and interface confidence metrics such as pLDDT, pTM, ipTM, ipAE, and ipSAE, as well as an ensemble-level interface for accessing and aggregating these metrics for a single input across multiple seeds and samples. We benchmark FoldKit on three types of AF3 co-folding datasets: (i) a protein design campaign with 2 chains per input, (ii) a TCR-pMHC dataset with 4 chains per input, and (iii) a pooled-AF3 protein-protein interaction dataset with up to 22 chains per input. We find that FoldKit reduces storage requirements by approximately 5-15-fold compared to native AF3 outputs, depending on dataset composition, while maintaining direct programmatic access to individual predictions, ensembles, and confidence metrics. By reducing storage requirements and facilitating programmatic access to relevant outputs, FoldKit facilitates large-scale computational studies of biomolecular interactions. FoldKit is available from PyPI and can be installed using pip.
Problem

Research questions and friction points this paper is trying to address.

co-folding
biomolecular complexes
data storage
large-scale analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

efficient storage
large-scale analysis
co-folding predictions
programmatic access
confidence metrics
🔎 Similar Papers
No similar papers found.
J
Jonathan A. Levine
Tri-Institutional Program in Computational Biology and Medicine, Weill Cornell Medicine, New York, NY USA; Laboratory of Lymphocyte Dynamics, Rockefeller University, New York, NY USA; The Halvorsen Center for Computational Oncology, Department of Epidemiology and Biostatistics, Memorial Sloan Kettering Cancer Center, New York, NY USA
M
Melissa Pathil
The Halvorsen Center for Computational Oncology, Department of Epidemiology and Biostatistics, Memorial Sloan Kettering Cancer Center, New York, NY USA; Computational and Systems Biology Program, Sloan Kettering Institute, Memorial Sloan Kettering Cancer Center, New York, NY, USA; Physiology, Biophysics & Systems Biology, Weill Cornell Medicine, Weill Cornell Medical College, New York, NY USA
S
Samuel Nitz
Tri-Institutional Program in Computational Biology and Medicine, Weill Cornell Medicine, New York, NY USA; Laboratory of Host-Pathogen Biology, The Rockefeller University, New York, NY 10065, USA
O
Olga Lyudovyk
The Halvorsen Center for Computational Oncology, Department of Epidemiology and Biostatistics, Memorial Sloan Kettering Cancer Center, New York, NY USA
B
Benjamin D. Greenbaum
The Halvorsen Center for Computational Oncology, Department of Epidemiology and Biostatistics, Memorial Sloan Kettering Cancer Center, New York, NY USA; Physiology, Biophysics & Systems Biology, Weill Cornell Medicine, Weill Cornell Medical College, New York, NY USA; The Olayan Center for Cancer Vaccines, Memorial Sloan Kettering Cancer Center, New York, NY USA