Multi-scale Decomposed Convolution Refinement Network for Visible-Infrared Person Re-Identification

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of significant cross-modality discrepancies and insufficient feature discriminability in visible-infrared person re-identification by proposing MDCRNet. The method introduces a hierarchical decomposition convolutional attention module and a granularity discriminative loss, integrating multi-scale decomposition convolutions, channel attention, and spatial awareness blocks. Furthermore, it employs a joint metric loss to simultaneously optimize intra-modal compactness and inter-modal separability, thereby significantly enhancing cross-modal feature learning and discrimination. Experimental results demonstrate that MDCRNet achieves state-of-the-art performance on both the SYSU-MM01 and RegDB datasets, effectively improving accuracy in cross-modality person re-identification tasks.
📝 Abstract
Visible-infrared person re-identification (VI-ReID) suffers from cross-modal discrepancies and limited discriminative capabilities, leading to suboptimal recognition performance. Current approaches exhibit limitations in semantic mining, cross-modal fusion and feature constraints. To tackle these challenges, we propose MDCRNet, a Multi-scale Decomposed Convolution Refinement Network that enhances cross-modal feature learning and discriminative metric learning. Specifically, we introduce a Hierarchical Learning Module (HLM) containing four Hierarchical Decomposed Convolution Attention (HDCA) modules, each equipped with lightweight channel attention and multi-scale spatial perception blocks to capture multi-scale spatial dependencies. Moreover, we develop a Joint Discriminative Metric Loss (JDML) incorporating a novel Granularity Discriminative Loss (GDL) that simultaneously optimizes intra-identity compactness and inter-identity separability across modalities. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate that MDCRNet achieves state-of-the-art performance on both benchmarks. Code is available at https://github.com/Kevin-zms/MDCRNet.
Problem

Research questions and friction points this paper is trying to address.

Visible-Infrared Person Re-Identification
Cross-modal discrepancies
Discriminative capabilities
Semantic mining
Cross-modal fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-scale Decomposed Convolution
Hierarchical Learning Module
Granularity Discriminative Loss
Visible-Infrared Person Re-Identification
Cross-modal Feature Learning
💼 Related Jobs
No related jobs found.
M
Mingsheng Zheng
School of Computer Science and Technology, Xinjiang University, Urumqi, Xinjiang 830046, China
Z
Zirui Jiang
School of Computer Science and Technology, Xinjiang University, Urumqi, Xinjiang 830046, China
B
Bo Liu
School of Computer Science and Technology, Xinjiang University, Urumqi, Xinjiang 830046, China
Y
Yupeng Chen
School of Computer Science and Technology, Xinjiang University, Urumqi, Xinjiang 830046, China
J
Jun Zhang
School of Computer Science and Technology, Xinjiang University, Urumqi, Xinjiang 830046, China
K
Kai Zhao
School of Computer Science and Technology, Xinjiang University, Urumqi, Xinjiang 830046, China; Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing, Xinjiang University, Urumqi, Xinjiang 830046, China