Benchmarking and Evaluation of AI Models in Biology: Outcomes and Recommendations from the CZI Virtual Cells Workshop
Biology lacks cross-domain, standardized AI model benchmarks, hindering model robustness and trustworthiness. To address this, we introduce the first multimodal AI benchmarking framework spanning imaging, transcriptomics, proteomics, and genomics—systematically tackling data heterogeneity, noise, bias, and resource fragmentation. Our approach integrates a high-fidelity data curation pipeline, unified preprocessing tools, biologically grounded multimodal evaluation metrics, and an open collaborative platform to enable fair, cross-task and cross-modal comparisons. A core innovation is the “virtual cell” paradigm—a biologically anchored, integrative evaluation framework—that unifies disparate modalities through shared cellular context. We further release a reproducible, extensible set of AI model evaluation guidelines. The framework significantly enhances rigor, transparency, and cross-domain comparability in biological AI research, accelerating AI-driven mechanistic discovery and therapeutic translation.