GraphT5: Unified Molecular Graph-Language Modeling via Multi-Modal Cross-Token Attention

📅 2025-03-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of cross-modal heterogeneity and semantic misalignment between SMILES strings (1D sequences) and molecular graphs (2D structures) in molecular language modeling. To this end, we propose a fine-grained cross-modal alignment framework built upon the T5 architecture. Our method jointly encodes SMILES tokens and graph-structured molecular representations—generated via a GNN—and introduces a learnable cross-token cross-attention mechanism that explicitly models semantic correspondences between SMILES tokens and graph nodes/edges. Unlike conventional approaches relying on simple concatenation or unimodal modeling, our mechanism enables deep integration of structural knowledge with sequence-based language generation. Empirically, the model achieves significant improvements over existing state-of-the-art methods on two key tasks: molecular caption generation and IUPAC name prediction. These results demonstrate that structure-aware multimodal co-modeling effectively enhances both molecular understanding and generative capabilities.

Technology Category

Application Category

📝 Abstract
Molecular language modeling tasks such as molecule captioning have been recognized for their potential to further understand molecular properties that can aid drug discovery or material synthesis based on chemical reactions. Unlike the common use of molecule graphs in predicting molecular properties, most methods in molecular language modeling rely heavily on SMILES sequences. This preference is because the task involves generating a sequence of multiple tokens using transformer-based models. Therefore, a main challenge is determining how to integrate graph data, which contains structural and spatial information about molecules, with text data. In addition, simply using both 1D SMILES text and 2D graph as inputs without addressing how they align and represent the molecule structure in different modalities makes it challenging to fully utilize structural knowledge about molecules. To this end, we propose GraphT5, a multi-modal framework that integrates 1D SMILES text and 2D graph representations of molecules for molecular language modeling. Specifically, we introduce a novel cross-token attention module in GraphT5 to bridge the gap arising from the fundamental differences between the two modalities of molecule representations. Cross-token attention exploits implicit information between SMILES and graphs of molecules, resulting from their interactions at a fine-grained token level that benefits molecular language modeling. Extensive experiments including molecule captioning, IUPAC name prediction tasks, and case studies show that our GraphT5 outperforms the latest baseline approaches, which validates the effectiveness of our GraphT5 in sufficiently utilizing 1D SMILES text and 2D graph representations.
Problem

Research questions and friction points this paper is trying to address.

Integrate 1D SMILES text and 2D graph representations for molecular language modeling.
Bridge gap between SMILES and graph modalities using cross-token attention.
Enhance molecular property understanding for drug discovery and material synthesis.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates 1D SMILES text and 2D graph representations
Introduces cross-token attention for modality alignment
Enhances molecular language modeling via multi-modal framework
💼 Related Jobs
No related jobs found.
S
Sangyeup Kim
Computer Science and Engineering, Seoul National University, South Korea
N
Nayeon Kim
Interdisciplinary Program in Artificial Intelligence, Seoul National University, South Korea
Y
Yinhua Piao
Computer Science and Engineering, Seoul National University, South Korea
S
Sun Kim
Interdisciplinary Program in Artificial Intelligence; Computer Science and Engineering, Seoul National University; AIGENDRUG Co., Ltd., South Korea