MODAL: Multi-Modal Object Re-ID via Model-Driven Sparse Decoupling and Text-Image Differential Filtering

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of feature entanglement, cross-modal conflict, and missing modalities in multi-modal re-identification by proposing the Modal framework. This approach innovatively integrates deep unfolding networks with coupled sparse coding to achieve transparent feature disentanglement, while incorporating a text-image differential filtering mechanism to enhance discriminability and robustness. Extensive experiments on four benchmark datasets demonstrate that the proposed framework achieves state-of-the-art performance, effectively mitigating performance degradation under missing modality scenarios. Furthermore, it significantly improves model interpretability, establishing a novel paradigm for multi-modal representation learning that simultaneously ensures transparency and robustness.
📝 Abstract
Multi-modal object re-identification (Re-ID) aims to facilitate cross-camera object retrieval in complex environments by leveraging complementary information from visual (e.g., RGB, NIR, TIR) and textual modalities. However, existing approaches often lack principled feature disentanglement and coherent multi-modal integration, leading to entangled representations that introduce cross-modal conflicts, obscure discriminative cues, and suffer distribution shift under modality-missing conditions. To tackle these challenges, we propose MODAL, a novel multi-modal object re-identification framework, grounded in coupled sparse coding theory and differential suppression principles. A core component of MODAL is a Multi-modal Feature Sparse Decoupling module, developed in a model-driven deep unrolling manner based on multi-modal coupled sparse coding. It explicitly decomposes multi-modal features into uni-modal specific, bi-modal and tri-modal shared representations, thereby achieving more transparent and effective feature disentanglement. Benefiting from the principled feature disentanglement, MODAL naturally mitigates performance degradation in incomplete-modality scenarios via a Modality-Aware Subspace Activation that selectively activates only the consistently shared subspaces. Moreover, we propose a Text-Image Differential Filtering module that leverages coarse-grained textual semantics to adaptively suppress task-irrelevant responses in the decoupled visual representations, thereby enhancing discriminative information. Extensive experiments on four datasets demonstrate that MODAL achieves state-of-the-art performance with superior transparency.
Problem

Research questions and friction points this paper is trying to address.

Multi-modal Object Re-ID
Feature Disentanglement
Cross-modal Conflict
Modality Missing
Distribution Shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model-Driven Sparse Decoupling
Text-Image Differential Filtering
Multi-Modal Object Re-ID
Coupled Sparse Coding
Modality-Aware Subspace Activation
C
Chengbo Huang
College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
Jun-Jie Huang
Jun-Jie Huang
National University of Defense Technology, Imperial College London
Inverse ProblemsSignal/Image ProcessingComputer VisionDeep Learning
L
Long Lan
College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
Tianrui Liu
Tianrui Liu
Associate Professor at National University of Defense Technology, PhD Imperial College London
Computer VisionMedical Image AnalysisDeep Learning
X
Xueqiong Li
College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
Y
Yuanxi Peng
College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
X
Xinwang Liu
College of Computer Science and Technology, National University of Defense Technology, Changsha 410073, China
M
Meng Wang
School of Computer Science and Information Engineering, Hefei University of Technology, Hefei 230002, China