🤖 AI Summary
Addressing the challenging problem of accurate registration between RGB and thermal infrared (RGB-T) images in intelligent transportation—exacerbated by large modality discrepancies—this paper proposes an end-to-end framework jointly modeling intra-modal self-correlation and cross-modal correspondence. We introduce a novel convolutional-Transformer hybrid architecture to separately capture intra-modal structural dependencies and cross-modal semantic consistency. A hierarchical optical flow decoder is further designed to progressively refine dense registration fields. Our method achieves significant improvements over state-of-the-art approaches on mainstream RGB-T benchmarks, demonstrating exceptional robustness under large disparities, severe occlusions, and adverse weather conditions. Moreover, it generalizes effectively to other cross-modal registration tasks—including RGB-near-infrared (RGB-N) and RGB-depth (RGB-D)—validating its strong cross-modal representation capability and broad applicability.
📝 Abstract
Multispectral imaging plays a critical role in a range of intelligent transportation applications, including advanced driver assistance systems (ADAS), traffic monitoring, and night vision. However, accurate visible and thermal (RGB-T) image registration poses a significant challenge due to the considerable modality differences. In this paper, we present a novel joint Self-Correlation and Cross-Correspondence Estimation Framework (SC3EF), leveraging both local representative features and global contextual cues to effectively generate RGB-T correspondences. For this purpose, we design a convolution-transformer-based pipeline to extract local representative features and encode global correlations of intra-modality for inter-modality correspondence estimation between unaligned visible and thermal images. After merging the local and global correspondence estimation results, we further employ a hierarchical optical flow estimation decoder to progressively refine the estimated dense correspondence maps. Extensive experiments demonstrate the effectiveness of our proposed method, outperforming the current state-of-the-art (SOTA) methods on representative RGB-T datasets. Furthermore, it also shows competitive generalization capabilities across challenging scenarios, including large parallax, severe occlusions, adverse weather, and other cross-modal datasets (e.g., RGB-N and RGB-D).