🤖 AI Summary
This work addresses the challenge posed by massive cell image datasets—such as JUMP (115 TB)—whose sheer scale and lack of standardized evaluation protocols hinder reproducible comparisons of representation methods. To overcome this, we introduce JUMP-lite, a high-quality 116 GB subset that reduces data volume by nearly three orders of magnitude while preserving phenotypic diversity through high-confidence perturbation filtering and lossy compression using JPEG XL. Leveraging the open-source framework Nahual, we integrate diverse representation approaches—including CellProfiler, MorphEM, OpenPhenom, SubCell, and DINOv2—for standardized benchmarking. Our evaluation demonstrates that compression does not compromise downstream performance and reveals significant differences among methods in terms of phenotypic activity and consistency, thereby enabling accessible, reproducible benchmarking for cellular image analysis.
📝 Abstract
Image-based profiling captures rich phenotypic signatures for drug discovery and functional genomics. Large public datasets like JUMP Cell Painting now provide millions of images for systematic study. However, the scale of these resources, 115 TB for JUMP alone, and fragmented evaluation practices make systematic comparison of representation methods intractable for many researchers. Here we present Nahual, an open-source framework for reproducible model deployment, and JUMP-lite, a curated 116 GB subset of JUMP that is 1000 times smaller while preserving phenotypic diversity through careful selection of perturbations with high-confidence annotations and a storage reduction via lossy JPEG XL compression. With these, we benchmark five representation methods, including classical features (CellProfiler) and deep learning models (MorphEM, OpenPhenom, SubCell, DINOv2), and demonstrate that compression preserves downstream signal while standardized phenotypic activity and consistency metrics reveal meaningful performance differences across methods. Together, JUMP-lite and Nahual provide a foundation for accessible, reproducible benchmarking of image-based cell representations.