Can Large Pretrained Depth Estimation Models Help With Image Dehazing?
Image dehazing faces three key challenges: significant spatial variation in haze distribution, poor generalization of existing methods, and a fundamental trade-off between accuracy and efficiency. To address these, this paper proposes a plug-and-play RGB-D fusion module. We systematically discover— for the first time—that depth features extracted from large-scale pre-trained depth estimation models exhibit strong consistency across multi-level haze conditions. Leveraging this inherent stability, we design a lightweight, architecture-agnostic feature fusion mechanism. The module integrates seamlessly into diverse mainstream dehazing networks without increasing inference overhead, thereby enhancing robustness and cross-scenario generalization. Extensive experiments on standard benchmarks—including SOTS and RESIDE—demonstrate substantial improvements in PSNR and SSIM, while maintaining real-time inference efficiency. Our approach thus bridges the gap between high-fidelity restoration and practical deployment requirements.