π€ AI Summary
This study addresses the high cost and low throughput of spatial transcriptomics, which hinder its application to large-scale tissue samples, by proposing VOICEβa multimodal foundation model that predicts single-cell gene expression from routine H&E images. Leveraging paired Xenium data, VOICE aligns histomorphological and gene expression embeddings through contrastive learning and employs a dual-branch architecture combining direct regression with reference-based retrieval. To account for varying morphological predictability across genes, the model incorporates gene-level adaptive weighting. Trained on 23 million cells, VOICE demonstrates strong generalization across held-out patients, tissue sections, and partially overlapping gene panels, consistently outperforming existing methods across seven evaluation metrics.
π Abstract
Spatial transcriptomics can resolve gene expression at single-cell resolution, but it is costly, limited to targeted panels of a few hundred to a few thousand genes, and applicable to only a small number of samples. H&E imaging, by contrast, is cheap and collected routinely at scale. This makes predicting single-cell expression directly from morphology a practical way to bring molecular analysis to large tissue archives. We therefore present VOICE, a multimodal foundation model that predicts single-cell gene expression from H&E images using paired Xenium data. VOICE first aligns cell centered H&E morphology from a pathology foundation model with single-cell expression embeddings from a transcriptome foundation model, trained using contrastive learning over 23 million cells. Next it predicts expression through two branches. One branch directly regresses expression from morphology. The other branch retrieves measured expression from similar reference cells, recovering genes that do not have morphological signal. Because genes vary in morphological predictability, VOICE fuses the two branches with a per-gene weight. After training, VOICE generalizes to heldout patients, slides, and partially overlapping gene panels from Xenium, and it consistently outperforms prior single-cell expression prediction methods on seven metrics.