CL-MVSNet: Unsupervised Multi-view Stereo with Dual-level Contrastive Learning
To address incomplete and brittle depth estimation in unsupervised multi-view stereo (MVS) caused by low-texture regions and view-dependent effects (e.g., reflections), this paper proposes a two-level contrastive learning framework: image-level and scene-level contrastive branches are jointly optimized to enhance contextual awareness and feature representation robustness. Additionally, we introduce an L₀.₅ photometric consistency loss that selectively emphasizes high-confidence correspondences, mitigating the over-penalization of low-gradient regions inherent in conventional L₁/L₂ losses. The method is fully unsupervised—requiring no ground-truth depth annotations. Evaluated on DTU and Tanks & Temples benchmarks, it achieves state-of-the-art performance among unsupervised MVS approaches and surpasses leading supervised methods without fine-tuning. Our core contributions are the first-ever dual-granularity contrastive mechanism for MVS and an L₀.₅ norm-driven photometric constraint, jointly advancing robustness and accuracy in texture-deficient and view-dependent scenarios.