HMGCLIP: Heterogeneous Multi-Granularity Contrastive Learning for E-commerce Representation Learning

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态大语言模型在精细属性区分上的不足,提出HMGCLIP框架,通过构建异构超图挖掘结构感知的难负样本,并在多个粒度上对齐语义。
📝 Abstract
Although recent Multimodal Large Language Models (MLLMs) have advanced general product understanding, they implicitly encode product information into global embeddings, thereby limiting their ability to capture fine-grained attributes. This limitation hinders performance in tasks requiring precise attribute discrimination, such as distinguishing subtle material differences among visually similar products. To address this challenge, we propose HMGCLIP, a unified multimodal embedding framework. By constructing a heterogeneous hypergraph, we leverage hypergraph topology to mine structure-aware hard negatives and align multi-granular semantics at both relation and hyperedge levels. This design enables a dual-granularity inference mechanism that dynamically fuses attribute evidence for both fine-grained and coarse-grained downstream tasks. Furthermore, we release a comprehensive fine-grained e-commerce dataset to facilitate future benchmarking. Extensive experiments on this new dataset and the public MAVE benchmark show that HMGCLIP outperforms strong multimodal encoders, MLLMs, and e-commerce baselines, validating the superiority of HMGCLIP.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
fine-grained attributes
product understanding
e-commerce
Innovation

Methods, ideas, or system contributions that make the work stand out.

Heterogeneous Hypergraph
Multi-Granularity Contrastive Learning
Structure-Aware Hard Negatives
Dual-Granularity Inference
Qiuyu Zhu
Qiuyu Zhu
Nanyang Technological University
Y
Yi Gao
Alibaba International Digital Commerce Group
Z
Zhichao Wan
Alibaba International Digital Commerce Group
M
Mingyang Ma
Alibaba International Digital Commerce Group