DMMRL: Disentangled Multi-Modal Representation Learning via Variational Autoencoders for Molecular Property Prediction
This work addresses the limitations of existing molecular property prediction methods, which often produce entangled representations that fail to disentangle structural, chemical, and functional factors and inadequately integrate multimodal information. To overcome these challenges, the authors propose a variational autoencoder-based disentangled representation learning framework that decomposes the molecular latent space into shared (structure-related) and private (modality-specific) subspaces. Orthogonality and alignment regularizations are introduced to enhance disentanglement, while a gated attention mechanism enables effective fusion of graph, sequence, and geometric modalities. Evaluated on seven benchmark datasets, the proposed method significantly outperforms current state-of-the-art models, achieving both improved predictive performance and enhanced interpretability of learned representations.