M2Depth: Unifying Monocular Depth Foundation Priors with Multi-View Stereo

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对多视图立体视觉在未见过场景中的泛化问题,提出了一种结合单目深度基础模型和级联MVS的双向互精炼框架,提高了深度估计的完整性和细节。
📝 Abstract
Deep learning-based Multi-View Stereo (MVS) has advanced significantly but often generalizes poorly to unseen scenes, particularly in occluded areas or regions with limited view overlap. To mitigate this, recent approaches integrate Depth Foundation Models (DFMs) into MVS pipelines to provide monocular depth priors. However, existing methods typically rely on a static, one-way fusion scheme, which fails to fully exploit the complementary strengths of both modalities. We propose a novel framework that overcomes this limitation by tightly coupling a DFM with a cascade MVS pipeline through a bidirectional mutual refinement strategy. Our method leverages MVS depth to resolve the scale ambiguity in monocular predictions, while the monocular depth, in turn, enhances the structural completeness and fine-grained detail of the MVS estimate. Furthermore, we introduce a prior-guided cost volume refinement mechanism that effectively integrates multi-view and monocular information via attention-based fusion and discretized depth bins, thereby promoting local geometric consistency. Extensive experiments demonstrate that our method outperforms state-of-the-art MVS approaches on standard benchmarks, producing more complete and generalizable depth maps with sharp boundaries. Furthermore, although not explicitly designed for sparse-view settings, our framework generalizes remarkably well, competing favorably with even dedicated sparse-view methods while maintaining a superior accuracy-efficiency trade-off.
Problem

Research questions and friction points this paper is trying to address.

Multi-View Stereo
Depth Foundation Models
Generalization
Occluded Areas
View Overlap
Innovation

Methods, ideas, or system contributions that make the work stand out.

bidirectional mutual refinement
prior-guided cost volume refinement
attention-based fusion
discretized depth bins
💼 Related Jobs
No related jobs found.
B
Byeonggwon Lee
Department of Computer Science and Artificial Intelligence, Dongguk University, Seoul 04620, Republic of Korea
S
Sanggi Lee
Department of Computer Science and Artificial Intelligence, Dongguk University, Seoul 04620, Republic of Korea
S
Siwoo Lee
Department of Computer Science and Artificial Intelligence, Dongguk University, Seoul 04620, Republic of Korea
K
Khang Truong Giang
42dot, Seongnam-si, Gyeonggi-do 13449, Republic of Korea
S
Soohwan Song
Department of Computer Science and Artificial Intelligence, Dongguk University, Seoul 04620, Republic of Korea