Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the performance degradation in visible-infrared object detection caused by severe cross-modal geometric misalignment. To this end, the authors propose JFRDet, an end-to-end network that introduces, for the first time, an explicit feature-domain affine alignment mechanism (CMAA) to correct multi-level feature misalignments. The method further incorporates an illumination-guided complementary fusion module (IGCF) and an alignment-quality consistency gating strategy (AQCG) to enable robust cross-modal feature fusion. Evaluated on the newly established DVMA benchmark, JFRDet achieves a state-of-the-art mAP50 of 69.7%, substantially improving detection robustness over existing approaches.
📝 Abstract
Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignment. Most existing methods rely on implicit feature adaptation to handle weakly misaligned scenarios, while large-offset geometric discrepancies remain insufficiently addressed. In this paper, we propose a Joint Feature-domain Registration and Detection network (JFRDet), an end-to-end visible-infrared oriented object detector tailored for severely cross-modal geometric discrepancies. JFRDet introduces a Cross-Modal Affine Alignment (CMAA) module to estimate an image-level affine transformation for explicit multi-level feature alignment. Note that illumination changes directly affect the reliability of RGB cues, an Illumination-Guided Complementary Fusion (IGCF) module adaptively exploits modality reliability under varying illumination conditions for cross-modal fusion. Then, an Alignment Quality-Consistency Gating (AQCG) strategy stabilizes joint optimization by modulating detection supervision according to alignment reliability and gradient consistency. We further construct DroneVehicle Misaligned (DVMA), a benchmark for evaluating visible-infrared oriented object detection under severe cross-modal geometric misalignment. The proposed JFRDet achieves 69.7\% $\mathrm{mAP}_{50}$ on DVMA, which represents state-of-the-art (SOTA) performance. The code and dataset will be available on GitHub.
Problem

Research questions and friction points this paper is trying to address.

cross-modal misalignment
visible-infrared object detection
geometric discrepancy
feature alignment
multimodal fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Modal Affine Alignment
Illumination-Guided Fusion
Feature-Domain Registration
Visible-Infrared Object Detection
Alignment Quality Consistency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Q
Qi Ming
Beijing University of Technology, China
Y
Yuyang Wang
Central South University, China
M
Mingjing Zhao
Beijing Electronics Science & Technology Institute
Y
Yifan Xiao
China Aerospace Science & Industry Corporation
Z
Zhixin Guo
China Aerospace Science & Industry Corporation
Zhiqiang Zhou
Zhiqiang Zhou
Beijing Institute of Technology
Computer VisionInformation Fusion
P
Peng Sun
Information Support Force Engineering University, China
J
Juan Fang
Beijing University of Technology, China
F
Fuqiang Yang
Trunk Technology (Beijing) Co., Ltd., China
Xudong Zhao
Xudong Zhao
Dalian University of Technology; Bohai University; China University of Petroleum
Switched systemsFuzzy systemsMarkovian jump systemsDelay systemsRobust control