ProtoHGF-Net: Prototype HyperGraph Fusion with Intra-modal Calibration for RGBT Object Detection
This work addresses the challenge in existing RGBT object detection methods, where dense cross-modal interactions are prone to background interference and struggle to focus on target semantics. To overcome this limitation, the authors propose a Prototype Hypergraph Fusion Network that reformulates cross-modal fusion as prototype-level semantic interaction. Additionally, they introduce a teacher-mask calibration distillation strategy to perform target-aware calibration of modality-specific features prior to fusion, effectively suppressing background noise and enhancing target representation. The proposed method achieves state-of-the-art performance with mAP50 scores of 85.9%, 88.2%, and 79.1% on the DroneVehicle, DVTOD, and FLIR datasets, respectively.