GENEB: Why Genomic Models Are Hard to Compare
This study addresses the lack of standardized evaluation protocols that hinder fair assessment of genomic foundation models’ performance and generalization. To this end, the authors introduce GENEB, a large-scale diagnostic benchmark that systematically evaluates frozen representations from 40 models across 100 tasks under a unified probing protocol, spanning 13 functional categories and supporting few-shot settings. This framework enables, for the first time, category-aware, fine-grained, and controllable multidimensional comparisons, revealing the instability of aggregate leaderboards and inherent trade-offs across tasks. Key findings indicate substantial variation in model rankings across functional categories, limited and inconsistent gains from increased model scale, and a more decisive influence of architectural design and alignment between pretraining data and downstream tasks than parameter count alone.