Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration
This work addresses the performance degradation in visible-infrared object detection caused by severe cross-modal geometric misalignment. To this end, the authors propose JFRDet, an end-to-end network that introduces, for the first time, an explicit feature-domain affine alignment mechanism (CMAA) to correct multi-level feature misalignments. The method further incorporates an illumination-guided complementary fusion module (IGCF) and an alignment-quality consistency gating strategy (AQCG) to enable robust cross-modal feature fusion. Evaluated on the newly established DVMA benchmark, JFRDet achieves a state-of-the-art mAP50 of 69.7%, substantially improving detection robustness over existing approaches.