Retrieval-guided Twin Fusion with Similarity-aware Contrast for Molecule-Text Alignment

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of fine-grained semantic alignment between molecular substructures and textual descriptions in molecule-text matching. To overcome this limitation, we propose RISEN, a novel framework that innovatively integrates a retrieval-guided twin fusion mechanism with similarity-aware soft-threshold contrastive learning. By constructing latent twin molecules to augment substructure semantic representations, RISEN achieves precise cross-modal fine-grained alignment. Extensive experiments demonstrate that RISEN significantly outperforms existing baselines across multiple benchmark datasets, effectively enhancing downstream tasks such as molecular search and property prediction. Consequently, this work establishes a new paradigm for cross-modal understanding in molecular science, bridging the gap between structural and textual modalities through enhanced semantic correspondence.
📝 Abstract
This paper studies the problem of molecule-text alignment, which aims to project molecules and their textual descriptions into a joint latent space for downstream tasks including molecule search and molecular property prediction. Previous approaches typically combine graph structure mining with contrastive learning to enhance joint representation learning. However, they typically neglect fine-grained semantic relationships between substructures and texts, leading to suboptimal performance on downstream tasks. Towards this end, we propose a novel approach named Retrieval-guided Twin Fusion with Similarity-aware Contrast (RISEN) for molecule-text alignment. The core idea of RISEN is to construct a latent twin molecule for each substructure with cross-modal retrieval for semantic enhancement. In particular, for each substructure query, we retrieve relevant textual descriptions and sample several molecules that share similar descriptions of substructures. Then, we aggregate their representations via attention pooling for a twin latent representation, which would be further fused with the original substructure for representation enrichment. In addition, we measure the similarity across substructures and texts, which would further guide cross-modal contrastive learning with soft thresholding. Extensive experiments on benchmark datasets validate the superiority of the proposed RISEN in comparison with existing baselines.
Problem

Research questions and friction points this paper is trying to address.

Molecule-Text Alignment
Fine-grained Semantic Relationships
Substructure-Text Matching
Joint Representation Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-guided Twin Fusion
Similarity-aware Contrast
Molecule-Text Alignment
Cross-modal Retrieval
Soft Thresholding
💼 Related Jobs
No related jobs found.
S
Shunshun Gu
University of Wisconsin–Madison, USA
S
Shengqi Qiu
University of Wisconsin–Madison, USA
H
Hang Zhou
University of North Carolina at Chapel Hill, USA
Xiao Luo
Xiao Luo
University of Wisconsin–Madison
Machine LearningLLMML for ScienceStatistical ModelingBioinformatics