🤖 AI Summary
This work addresses critical limitations in existing anticancer drug response prediction models—namely, small data scales, narrow coverage of cancer types and compounds, and the absence of a unified benchmark hindering reliable model comparison. By integrating multi-source pharmacogenomic data from resources such as PharmacoDB, the study substantially expands the IMPROVE benchmark to create the largest and most chemically diverse AI-ready dataset to date, encompassing millions of drug response measurements, extensive multi-omics features, and over 50,000 novel compounds. Through standardized data schemas, multi-omics integration, and rigorous evaluation protocols—including drug-blind, cancer-blind, and disjoint splits—the expanded benchmark demonstrates significantly improved model generalization to unseen compounds, consistently outperforming the original benchmark across key test scenarios.
📝 Abstract
Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective anticancer drugs. However, their predictive performance is often constrained by limited dataset scale and insufficient coverages of cancer and chemical spaces. In addition, inconsistent benchmarking practices hinder reliable comparison across models. Standardized frameworks, such as the Innovative Methodologies and New Data for Predictive Oncology Model Evaluation (IMPROVE) project, provide unified data schemas and evaluation protocols for consistent benchmarking, but improving model generalizability requires larger and more diverse training data. In this work, we substantially expand the IMPROVE benchmark through large-scale integration of pharmacogenomic data, primarily from PharmacoDB, together with additional smaller data sources. The expanded resource includes millions of drug response measurements, broader multi-omics coverage, and a major increase in chemical diversity, adding more than 50,000 compounds. To evaluate the impact of the new dataset compared to the original IMPROVE benchmark dataset, we trained DRP models using the two datasets and assess their prediction performance using a common test set and several evaluation strategies, including drug-blind, cancer-blind, and disjoint data splits. While cancer-blind performance remained comparable to the original benchmark, models trained on the expanded dataset showed consistent improvements in drug-blind and disjoint settings, indicating enhanced generalization to previously unseen compounds. These results position the expanded dataset as a community resource that provides a richer foundation for developing DRP models intended to aid in the discovery of novel anticancer drugs.