MyGram: Modality-aware Graph Transformer with Global Distribution for Multi-modal Entity Alignment
This work addresses the limitations of existing multimodal entity alignment methods, which are often susceptible to superficial feature interference and overlook deep structural context within individual modalities. To overcome these issues, the authors propose MyGram, a novel model that integrates image and text semantics through a modality-aware graph Transformer and captures high-order intra-modal structural information via a modality diffusion learning module. Furthermore, they introduce a Gram Loss regularizer based on minimizing the volume of a four-dimensional parallelotope to enforce global cross-modal distributional consistency. Extensive experiments on five benchmark datasets demonstrate that MyGram significantly outperforms state-of-the-art approaches, achieving Hits@1 improvements of 4.8%, 9.9%, and 4.3% on FBDB15K, FBYG15K, and DBP15K, respectively.