Institution profile

Guangdong Industry Polytechnic University

Academic institutionasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

MyGram: Modality-aware Graph Transformer with Global Distribution for Multi-modal Entity Alignment

Jan 17, 2026

This work addresses the limitations of existing multimodal entity alignment methods, which are often susceptible to superficial feature interference and overlook deep structural context within individual modalities. To overcome these issues, the authors propose MyGram, a novel model that integrates image and text semantics through a modality-aware graph Transformer and captures high-order intra-modal structural information via a modality diffusion learning module. Furthermore, they introduce a Gram Loss regularizer based on minimizing the volume of a four-dimensional parallelotope to enforce global cross-modal distributional consistency. Extensive experiments on five benchmark datasets demonstrate that MyGram significantly outperforms state-of-the-art approaches, achieving Hits@1 improvements of 4.8%, 9.9%, and 4.3% on FBDB15K, FBYG15K, and DBP15K, respectively.

0 citationsRead paper

MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering

Jan 05, 2026arXiv.org

This work addresses the challenge in continual visual question answering (Continual VQA) of simultaneously preserving prior knowledge, adapting to new information, and maintaining robust feature representations. To this end, the authors propose a synergistic mechanism that integrates multimodal information, prototype-based memory management, and global noise filtering. The approach employs adaptive memory allocation to dynamically optimize knowledge storage and suppresses cross-task interference at the feature level, thereby enabling efficient knowledge acquisition, retention, and compositional generalization. Evaluated across 10 continual VQA tasks, the model achieves average accuracies of 43.38% on standard tasks and 42.53% on novel compositional tasks, with remarkably low average forgetting rates of 2.32% and 3.60%, respectively, substantially outperforming existing methods.

0 citationsRead paper

KeenKT: Knowledge Mastery-State Disambiguation for Knowledge Tracing

Dec 21, 2025

In knowledge tracing, point estimation fails to disentangle students’ true proficiency from behavioral noise, resulting in ambiguous mastery state modeling. To address this, we propose the first NIG-KT framework, which models latent knowledge states at each interaction step using the Normal-Inverse Gaussian (NIG) distribution—explicitly decoupling ability estimation from epistemic and aleatoric uncertainty. We further introduce an NIG-distance-based attention mechanism to enhance sensitivity to learning dynamics and transient performance fluctuations. Additionally, we design a diffusion-driven denoising reconstruction objective coupled with distributional contrastive learning to jointly optimize robust state estimation and uncertainty-aware representation. Extensive experiments on six benchmark datasets demonstrate consistent superiority over state-of-the-art methods: AUC improves by up to 5.85% and ACC by up to 6.89%, significantly enhancing both predictive accuracy and robustness to behavioral noise.

0 citationsRead paper

Structurally Refined Graph Transformer for Multimodal Recommendation

Nov 01, 2025

To address three key challenges in multimodal recommendation—redundant information interference, disconnection between local and global semantics, and insufficient modeling of user-item interactions—this paper proposes SRGFormer, a novel multimodal recommendation model integrating hypergraph neural networks with Transformer architectures. Its core contributions are: (1) constructing a multimodal hypergraph structure to explicitly capture high-order user-item interactions; (2) designing a dual-path semantic encoder that jointly learns local fine-grained features and global collaborative patterns; and (3) introducing cross-modal contrastive self-supervised learning to enhance preference discrimination under sparse interaction scenarios. Extensive experiments on the Sports, Electronics, and Clothing datasets demonstrate that SRGFormer consistently outperforms state-of-the-art methods, achieving an average improvement of 4.47% on the Sports dataset and significantly boosting purchase behavior prediction accuracy. The source code is publicly available.

0 citationsRead paper
Recent publications

Latest Papers

MyGram: Modality-aware Graph Transformer with Global Distribution for Multi-modal Entity Alignment

Jan 17, 2026

This work addresses the limitations of existing multimodal entity alignment methods, which are often susceptible to superficial feature interference and overlook deep structural context within individual modalities. To overcome these issues, the authors propose MyGram, a novel model that integrates image and text semantics through a modality-aware graph Transformer and captures high-order intra-modal structural information via a modality diffusion learning module. Furthermore, they introduce a Gram Loss regularizer based on minimizing the volume of a four-dimensional parallelotope to enforce global cross-modal distributional consistency. Extensive experiments on five benchmark datasets demonstrate that MyGram significantly outperforms state-of-the-art approaches, achieving Hits@1 improvements of 4.8%, 9.9%, and 4.3% on FBDB15K, FBYG15K, and DBP15K, respectively.

0 citationsRead paper

MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering

Jan 05, 2026arXiv.org

This work addresses the challenge in continual visual question answering (Continual VQA) of simultaneously preserving prior knowledge, adapting to new information, and maintaining robust feature representations. To this end, the authors propose a synergistic mechanism that integrates multimodal information, prototype-based memory management, and global noise filtering. The approach employs adaptive memory allocation to dynamically optimize knowledge storage and suppresses cross-task interference at the feature level, thereby enabling efficient knowledge acquisition, retention, and compositional generalization. Evaluated across 10 continual VQA tasks, the model achieves average accuracies of 43.38% on standard tasks and 42.53% on novel compositional tasks, with remarkably low average forgetting rates of 2.32% and 3.60%, respectively, substantially outperforming existing methods.

0 citationsRead paper

KeenKT: Knowledge Mastery-State Disambiguation for Knowledge Tracing

Dec 21, 2025

In knowledge tracing, point estimation fails to disentangle students’ true proficiency from behavioral noise, resulting in ambiguous mastery state modeling. To address this, we propose the first NIG-KT framework, which models latent knowledge states at each interaction step using the Normal-Inverse Gaussian (NIG) distribution—explicitly decoupling ability estimation from epistemic and aleatoric uncertainty. We further introduce an NIG-distance-based attention mechanism to enhance sensitivity to learning dynamics and transient performance fluctuations. Additionally, we design a diffusion-driven denoising reconstruction objective coupled with distributional contrastive learning to jointly optimize robust state estimation and uncertainty-aware representation. Extensive experiments on six benchmark datasets demonstrate consistent superiority over state-of-the-art methods: AUC improves by up to 5.85% and ACC by up to 6.89%, significantly enhancing both predictive accuracy and robustness to behavioral noise.

0 citationsRead paper

Structurally Refined Graph Transformer for Multimodal Recommendation

Nov 01, 2025

To address three key challenges in multimodal recommendation—redundant information interference, disconnection between local and global semantics, and insufficient modeling of user-item interactions—this paper proposes SRGFormer, a novel multimodal recommendation model integrating hypergraph neural networks with Transformer architectures. Its core contributions are: (1) constructing a multimodal hypergraph structure to explicitly capture high-order user-item interactions; (2) designing a dual-path semantic encoder that jointly learns local fine-grained features and global collaborative patterns; and (3) introducing cross-modal contrastive self-supervised learning to enhance preference discrimination under sparse interaction scenarios. Extensive experiments on the Sports, Electronics, and Clothing datasets demonstrate that SRGFormer consistently outperforms state-of-the-art methods, achieving an average improvement of 4.47% on the Sports dataset and significantly boosting purchase behavior prediction accuracy. The source code is publicly available.

0 citationsRead paper