🤖 AI Summary
Current evaluations of pretrained molecular embedding models lack statistical rigor and fair benchmarking, often assuming neural models inherently outperform traditional fingerprints like ECFP without empirical validation. Method: We systematically evaluate 25 pretrained molecular embedding models across 25 benchmark datasets and introduce a hierarchical Bayesian statistical testing framework to control for multiple hypothesis testing bias in a unified manner. Contribution/Results: Contrary to prevailing assumptions, most pretrained models fail to achieve statistically significant improvements over the ECFP baseline; only CLAMP demonstrates robust superiority. This challenges the implicit “neural models are necessarily better” assumption pervasive in molecular representation learning and exposes systemic issues—including methodologically unsound evaluation protocols, weak baselines, and reporting bias. Based on these findings, we propose a standardized evaluation protocol, recommend strong baselines (e.g., ECFP with optimized hyperparameters), and advocate for small-sample validation to ensure reliability. Our work establishes a methodological foundation and practical guidelines for trustworthy molecular AI research.
📝 Abstract
Pretrained neural networks have attracted significant interest in chemistry and small molecule drug design. Embeddings from these models are widely used for molecular property prediction, virtual screening, and small data learning in molecular chemistry. This study presents the most extensive comparison of such models to date, evaluating 25 models across 25 datasets. Under a fair comparison framework, we assess models spanning various modalities, architectures, and pretraining strategies. Using a dedicated hierarchical Bayesian statistical testing model, we arrive at a surprising result: nearly all neural models show negligible or no improvement over the baseline ECFP molecular fingerprint. Only the CLAMP model, which is also based on molecular fingerprints, performs statistically significantly better than the alternatives. These findings raise concerns about the evaluation rigor in existing studies. We discuss potential causes, propose solutions, and offer practical recommendations.