Disentangled Shared Representations Improve Morpho-Transcriptomic Integration

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the coupling of shared and modality-specific variations in multimodal spatial transcriptomics representations by proposing an explicit disentanglement framework. The approach separates shared and private latent components between H&E histology and spatial transcriptomics data to optimize representation learning. A systematic evaluation benchmarking VAEs, contrastive learning, and their disentangled variants on cross-modal reconstruction and probing tasks demonstrates that contrastive objectives outperform traditional VAEs, while disentanglement mechanisms significantly enhance performance across all metrics. These findings validate the efficacy of explicit disentanglement for multimodal representation learning, establishing a superior paradigm for downstream spatial transcriptomics analyses.
📝 Abstract
Spatial transcriptomics (ST) enables the simultaneous profiling of gene expression and tissue morphology, creating an opportunity to learn multimodal representations capturing shared morpho-transcriptomic structure. However, standard multimodal models often compress modalities into a common latent space without explicitly separating shared and modality-specific sources of variation, which may limit downstream utility. We investigate whether explicit disentanglement of shared and private latent components improves multimodal representation learning for paired Hematoxylin \& Eosin (H\&E) and ST data. We compare VAE-based and contrastive approaches, each in standard and disentangled variants, across two cancer cohorts under matched experimental conditions. Representations are evaluated using cross-modal reconstruction, downstream probing and cross-modal probe transfer. The experiments suggest two main trends. First, contrastive objectives yield higher downstream probing performance than VAE-based models. Second, disentangled variants improve the selected reconstruction and probing metrics, although the gains depend on the model family, task, direction, and disentanglement strength. Overall, our results suggest that explicitly factorizing shared and modality-specific information can improve multimodal representation learning for spatial transcriptomics and provides a useful evaluation framework for future foundation models.
Problem

Research questions and friction points this paper is trying to address.

Spatial Transcriptomics
Multimodal Representation Learning
Disentangled Representations
Morpho-Transcriptomic Integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Disentangled Representation
Spatial Transcriptomics
Multimodal Integration
Contrastive Learning
Morpho-Transcriptomic
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Julian Ostermaier
ESPCI Paris, PSL University, Paris, France
S
Swann Ruyter
Sorbonne Université, CNRS, Inserm, AP-HP, Inria, Paris Brain Institute – ICM, Paris, France
Reuben Dorent
Reuben Dorent
Inria
Machine LearningDeep LearningMedical Image Analysis
D
Daniel Racoceanu
Sorbonne Université, CNRS, Inserm, AP-HP, Inria, Paris Brain Institute – ICM, Paris, France