VOICE: A Vision-Omics Foundation Model Integrating Direct and Retrieval-Based Prediction of In-situ Single-Cell Gene Expression

πŸ“… 2026-08-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high cost and low throughput of spatial transcriptomics, which hinder its application to large-scale tissue samples, by proposing VOICEβ€”a multimodal foundation model that predicts single-cell gene expression from routine H&E images. Leveraging paired Xenium data, VOICE aligns histomorphological and gene expression embeddings through contrastive learning and employs a dual-branch architecture combining direct regression with reference-based retrieval. To account for varying morphological predictability across genes, the model incorporates gene-level adaptive weighting. Trained on 23 million cells, VOICE demonstrates strong generalization across held-out patients, tissue sections, and partially overlapping gene panels, consistently outperforming existing methods across seven evaluation metrics.
πŸ“ Abstract
Spatial transcriptomics can resolve gene expression at single-cell resolution, but it is costly, limited to targeted panels of a few hundred to a few thousand genes, and applicable to only a small number of samples. H&E imaging, by contrast, is cheap and collected routinely at scale. This makes predicting single-cell expression directly from morphology a practical way to bring molecular analysis to large tissue archives. We therefore present VOICE, a multimodal foundation model that predicts single-cell gene expression from H&E images using paired Xenium data. VOICE first aligns cell centered H&E morphology from a pathology foundation model with single-cell expression embeddings from a transcriptome foundation model, trained using contrastive learning over 23 million cells. Next it predicts expression through two branches. One branch directly regresses expression from morphology. The other branch retrieves measured expression from similar reference cells, recovering genes that do not have morphological signal. Because genes vary in morphological predictability, VOICE fuses the two branches with a per-gene weight. After training, VOICE generalizes to heldout patients, slides, and partially overlapping gene panels from Xenium, and it consistently outperforms prior single-cell expression prediction methods on seven metrics.
Problem

Research questions and friction points this paper is trying to address.

spatial transcriptomics
single-cell gene expression
H&E imaging
morphology-based prediction
gene expression prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal foundation model
spatial transcriptomics
H&E image
contrastive learning
retrieval-based prediction
πŸ”Ž Similar Papers
No similar papers found.