Structurally Refined Graph Transformer for Multimodal Recommendation

📅 2025-11-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address three key challenges in multimodal recommendation—redundant information interference, disconnection between local and global semantics, and insufficient modeling of user-item interactions—this paper proposes SRGFormer, a novel multimodal recommendation model integrating hypergraph neural networks with Transformer architectures. Its core contributions are: (1) constructing a multimodal hypergraph structure to explicitly capture high-order user-item interactions; (2) designing a dual-path semantic encoder that jointly learns local fine-grained features and global collaborative patterns; and (3) introducing cross-modal contrastive self-supervised learning to enhance preference discrimination under sparse interaction scenarios. Extensive experiments on the Sports, Electronics, and Clothing datasets demonstrate that SRGFormer consistently outperforms state-of-the-art methods, achieving an average improvement of 4.47% on the Sports dataset and significantly boosting purchase behavior prediction accuracy. The source code is publicly available.

Technology Category

Application Category

📝 Abstract
Multimodal recommendation systems utilize various types of information, including images and text, to enhance the effectiveness of recommendations. The key challenge is predicting user purchasing behavior from the available data. Current recommendation models prioritize extracting multimodal information while neglecting the distinction between redundant and valuable data. They also rely heavily on a single semantic framework (e.g., local or global semantics), resulting in an incomplete or biased representation of user preferences, particularly those less expressed in prior interactions. Furthermore, these approaches fail to capture the complex interactions between users and items, limiting the model's ability to meet diverse users. To address these challenges, we present SRGFormer, a structurally optimized multimodal recommendation model. By modifying the transformer for better integration into our model, we capture the overall behavior patterns of users. Then, we enhance structural information by embedding multimodal information into a hypergraph structure to aid in learning the local structures between users and items. Meanwhile, applying self-supervised tasks to user-item collaborative signals enhances the integration of multimodal information, thereby revealing the representational features inherent to the data's modality. Extensive experiments on three public datasets reveal that SRGFormer surpasses previous benchmark models, achieving an average performance improvement of 4.47 percent on the Sports dataset. The code is publicly available online.
Problem

Research questions and friction points this paper is trying to address.

Distinguishing redundant from valuable multimodal recommendation data
Overcoming single semantic framework limitations in user preference modeling
Capturing complex user-item interactions for diverse recommendation needs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Modified transformer captures overall user behavior patterns
Embedded multimodal data into hypergraph for local structures
Applied self-supervised tasks to enhance multimodal integration
🔎 Similar Papers
No similar papers found.
K
Ke Shi
School of Computer Science, Hubei University, Wuhan 430062, China, Hubei Key Laboratory of Big Data Intelligent Analysis and Application, Hubei University, Wuhan 430062, China, and Key Laboratory of Intelligent Sensing System and Security (Hubei University), Ministry of Education, Wuhan 430062, China
Y
Yan Zhang
School of Computer Science, Hubei University, Wuhan 430062, China, Hubei Key Laboratory of Big Data Intelligent Analysis and Application, Hubei University, Wuhan 430062, China, and Key Laboratory of Intelligent Sensing System and Security (Hubei University), Ministry of Education, Wuhan 430062, China
M
Miao Zhang
School of Computer Science, Hubei University, Wuhan 430062, China, Hubei Key Laboratory of Big Data Intelligent Analysis and Application, Hubei University, Wuhan 430062, China, and Key Laboratory of Intelligent Sensing System and Security (Hubei University), Ministry of Education, Wuhan 430062, China
L
Lifan Chen
School of Computer Science, Hubei University, Wuhan 430062, China, Hubei Key Laboratory of Big Data Intelligent Analysis and Application, Hubei University, Wuhan 430062, China, and Key Laboratory of Intelligent Sensing System and Security (Hubei University), Ministry of Education, Wuhan 430062, China
J
Jiali Yi
School of Computer Science, Hubei University, Wuhan 430062, China, Hubei Key Laboratory of Big Data Intelligent Analysis and Application, Hubei University, Wuhan 430062, China, and Key Laboratory of Intelligent Sensing System and Security (Hubei University), Ministry of Education, Wuhan 430062, China
K
Kui Xiao
School of Computer Science, Hubei University, Wuhan 430062, China, Hubei Key Laboratory of Big Data Intelligent Analysis and Application, Hubei University, Wuhan 430062, China, and Key Laboratory of Intelligent Sensing System and Security (Hubei University), Ministry of Education, Wuhan 430062, China
X
Xiaoju Hou
Institute of Vocational Education, Guangdong Industry Polytechnic University, Guangzhou 510300, China
Zhifei Li
Zhifei Li
Research Scientist at Google
machine translationnatural language processingmachine learningwireless networks