P2Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the distortion in infrared and visible image fusion caused by inherent modality discrepancies and the limitations of existing methods that rely on static constraints or external priors while neglecting intrinsic modality characteristics. To this end, we propose a dual-intrinsic-prompt distillation framework that transforms thermal saliency and spatial quality priors into learnable dynamic prompts. Our approach enables decoupled and adaptive modality-specific feature fusion through a Teach-to-Fuse dual-granularity progressive guidance mechanism and a Gated Dynamic Expert Recalibration (GDER) module. The method achieves state-of-the-art performance across five benchmark datasets, outperforming competitors in 14 out of 20 evaluation metrics, and significantly enhances downstream object detection—improving mAP by 3.2%, 0.5%, and 0.9% on MSRS, M3FD, and DroneVehicle, respectively.
📝 Abstract
Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundamental challenge. Existing prior-guided methods often rely on static constraints that induce optimization conflicts or utilize extrinsic semantic priors from large-scale foundation models (e.g., CLIP/DINO), which frequently fail to exploit the intrinsic modality characteristics essential for high-fidelity fusion. To address these issues, we propose P2Fusion, a prior-guided distillation-based framework that reformulates IVIF via dual intrinsic prompts. Instead of imposing hard-coded penalties, we distill image-intrinsic priors, thermal saliency and spatial quality, into learnable dynamic regulators. Specifically, a Teach-to-Fuse mechanism provides dual-granularity progressive guidance, coupled with a Gated Dynamic Expert Recalibration (GDER) module for decoupled feature refinement. This design enables the network to adaptively mediate modal competition through expert specialization. Extensive experiments demonstrate that P2Fusion achieves state-of-the-art performance across five mainstream datasets. Notably, our framework demonstrates consistent performance advantages in fusion quality, achieving state-of-the-art results in 14 out of 20 key evaluation metrics across 5 benchmarks. Furthermore, it effectively contributes to the robustness of downstream perception, such as +3.2% mAP on MSRS, +0.5% mAP on M3FD and +0.9% mAP on DroneVehicle for object detection. Our code will be available at https://github.com/YiShi99/P2Fusion
Problem

Research questions and friction points this paper is trying to address.

Infrared-visible image fusion
information disparity
modality characteristics
prior-guided fusion
multimodal perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

dual-prior distillation
prompt-based fusion
dynamic expert recalibration
intrinsic modality priors
progressive guidance
💼 Related Jobs
No related jobs found.
Y
Yi Shi
Northwestern Polytechnical University, Xi’an, China
H
Huichao Xie
Northwestern Polytechnical University, Xi’an, China
Y
Yuqing Wang
Northwestern Polytechnical University, Xi’an, China
M
Mingyu Wang
Northwestern Polytechnical University, Xi’an, China
K
Kaihui Yang
Northwestern Polytechnical University, Xi’an, China
Yu Liu
Yu Liu
Professor, Hefei University of Technology
Image FusionImage RestorationMedical Image ProcessingComputer Vision
R
Ruitao Lu
Rocket Force University of Engineering, Xi’an, China
L
Lizhe Li
Northwestern Polytechnical University, Xi’an, China
J
Junwei Han
Northwestern Polytechnical University, Xi’an, China; Chongqing University of Posts and Telecommunications, Chongqing, China
D
Dingwen Zhang
Northwestern Polytechnical University, Xi’an, China