Same Encoder, Different Winner: A Paired-View Framework for Cell Painting Encoder Evaluation

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入CP-BG-Bench框架,使用四种匹配视图评估细胞图像编码器在不同协议下的表现差异,揭示了单一评价指标的局限性。
📝 Abstract
Vision encoders for Cell Painting are typically ranked by a single evaluation, commonly replicate mean average precision (mAP). We introduce CP-BG-Bench, a paired-view evaluation framework that holds the central cell fixed across four matched views (raw crop C, segmented S, and density-augmented variants CD and SD), ablating or augmenting surrounding pixels as a controlled intervention. Instantiating the framework on three datasets (JUMP-CP, RxRx1, RxRx3-core) and three encoders (DINOv3 ViT-B/16, OpenPhenom, SubCell) under four community-standard protocols (replicate mAP, scIB batch integration, CellProfiler feature prediction, cross-batch perturbation recall), we find that the four protocols rank the same encoders systematically differently, with disagreements decomposing along three axes: cell versus background, morphology versus context, and within-study versus across-batch. The largest effect: on RxRx3-core, SubCell with segmented inputs retains 94% of crop replicate mAP but only 32% of crop R@10, so the within-study signal preserved under segmentation is largely non-transferable; density augmentation recovers 84% of the within-study C-to-S gap but only 8% of the cross-batch gap. Segmented views predict CellProfiler features as well as or better than crops on two of three datasets, inverting the replicate-mAP ranking, and the C-to-S gap varies by an order of magnitude across datasets while remaining similar across encoders, indicating that background-driven gain is set by experimental design rather than by the encoder. Single-metric ranking of Cell Painting encoders is therefore sensitive to the protocol used, and protocol disagreements are interpretable as projections onto the three axes the paired-view design exposes. We will release the paired-view datasets, reconstruction pipelines, 36 trained checkpoints, aggregated embeddings, and the full evaluation suite.
Problem

Research questions and friction points this paper is trying to address.

Cell Painting
Encoder Evaluation
Replicate mAP
Paired-View Framework
Background-driven Gain
Innovation

Methods, ideas, or system contributions that make the work stand out.

paired-view framework
cell painting encoder evaluation
density augmentation
protocol disagreement
background-driven gain
🔎 Similar Papers
T
Tim Treis
Institute of Computational Biology, Helmholtz Munich
N
Nikita Moshkov
Institute of Computational Biology, Helmholtz Munich
Johan Fredin Haslum
Johan Fredin Haslum
PhD Student, Machine Learning, KTH - Royal Institute of Technology
Shantanu Singh
Shantanu Singh
Broad Institute of MIT and Harvard
F
Fabian J. Theis
Institute of Computational Biology, Helmholtz Munich