Reading the Whole Heart: Latent-Attention Masked Autoencoders for Multimodal Cardiac Representation Learning

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过引入Latent-Attention Masked Autoencoders (LAMAE)解决心血管诊断中多模态数据整合问题,利用自监督预训练学习患者级表示,优于传统方法。
📝 Abstract
Cardiovascular diagnosis rests on integrating complementary modalities, like ECG, echocardiography, chest radiographs, and clinical variables, each capturing distinct but correlated aspects of cardiac physiology. Yet most medical foundation models remain modality-specific, combining modalities only for finetuning or post-training. This discards the cross-modal evidence clinicians naturally integrate and ignores the structure within each modality. We introduce Latent-Attention Masked Autoencoders (LAMAE), a multimodal, structure-aware masked autoencoder that jointly learns patient-level representations during self-supervised pretraining. Rather than fusing modalities post hoc, LAMAE exchanges information directly in the latent space through a shared latent-attention module operating over a study-view-entity hierarchy, enabling aggregation of variable observations and graceful handling of missing modalities. Pretrained on over 1.2 million MIMIC-IV hospital stays, LAMAE outperforms modality-specific pretraining and strong contrastive and vision-language baselines across multimodal hospital-stay tasks, such as in-hospital mortality, ICD-10 and DRG coding, and length of stay, while remaining competitive on unimodal tasks. These gains persist even when only a single modality is available at test time, showing that modeling both intra- and inter-modal structure yields more robust, transferable representations.
Problem

Research questions and friction points this paper is trying to address.

multimodal cardiac representation
cross-modal evidence
structure-aware
latent space
missing modalities
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent-Attention Masked Autoencoders
multimodal representation learning
structure-aware
latent space
shared latent-attention module
A
Andrea Agostini
Department of Computer Science, ETH Zurich, Switzerland
S
Simon Böhi
Department of Biomedical Engineering, University of Basel, Switzerland
Moritz Vandenhirtz
Moritz Vandenhirtz
PhD student, ETH Zurich
Generative ModelingInterpretable Machine LearningComputer VisionMedical Data Science
S
Samuel Ruiperez-Campillo
Department of Computer Science, ETH Zurich, Switzerland
M
Max Krähenmann
Department of Biomedical Engineering, University of Basel, Switzerland
S
Silke Mühlstedt
Department of Computer Science, ETH Zurich, Switzerland
Irene Cannistraci
Irene Cannistraci
PostDoctoral Researcher, ETH Zurich
Deep LearningRepresentation Learning
E
Ece Özkan Elsen
Department of Biomedical Engineering, University of Basel, Switzerland
J
Julia E. Vogt
Department of Computer Science, ETH Zurich, Switzerland
Thomas M. Sutter
Thomas M. Sutter
Postdoc, ETH Zurich
Generative ModelsMultimodal MLProbabilistic MLRepresentation LearningML for Healthcare