Efficient and High-Quality Depth Estimation via Pixel-Space Diffusion with Linear Attention

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究提出Lapis框架,通过线性注意力和一步扩散方法解决高分辨率图像深度估计中的计算成本和细节损失问题。
📝 Abstract
This work presents $\textbf{Lapis}$, a $\textbf{l}$inear-$\textbf{a}$ttention-based $\textbf{pi}$xel-$\textbf{s}$pace generative framework that achieves efficient and high-fidelity depth estimation with one-step diffusion. While generative frameworks have significantly advanced monocular depth estimation with superior detail fidelity, the $\mathcal{O}(N^2)$ complexity of standard attention and the multi-step denoising process introduce prohibitive computational costs when scaling them to high-resolution image applications. Although linear attention and one-step prediction are intuitively viable, directly applying them leads to poor structural consistency, detail loss, and noise. Lapis rectifies these limitations through a coarse-to-fine hierarchy. Specifically, a Patch-level Consistency Module restores structural coherence by integrating semantic and spatial priors. Subsequently, a Pixel-level Refinement Module recovers sharp geometric boundaries via skip-connection-based pixel correspondence. Furthermore, to mitigate sampling noise inherent in one-step diffusion, we leverage the manifold assumption and adopt a direct $\mathbf{x}$-prediction strategy to target the clean data manifold. Extensive evaluations on multiple benchmarks demonstrate that Lapis consistently achieves state-of-the-art (SOTA) accuracy and boundary sharpness across various resolutions, reducing inference latency by up to 7.6$\times$ at 1080P and 10.9$\times$ at 1440P resolution compared to previous SOTA generative models.
Problem

Research questions and friction points this paper is trying to address.

Depth Estimation
Linear Attention
One-step Diffusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Linear Attention
One-Step Diffusion
Coarse-to-Fine Hierarchy
Patch-Level Consistency
Pixel-Level Refinement
💼 Related Jobs
No related jobs found.
B
Bingde Liu
MoE Key Lab of Artificial Intelligence, Institute of AI, Shanghai Jiao Tong University, Shanghai, China; Zhiyuan College, Shanghai Jiao Tong University, Shanghai, China
W
Wu Ran
MoE Key Lab of Artificial Intelligence, Institute of AI, Shanghai Jiao Tong University, Shanghai, China
Jinglei Zhang
Jinglei Zhang
Shanghai Jiao Tong University
computer vision
H
Huanhuan Yuan
MoE Key Lab of Artificial Intelligence, Institute of AI, Shanghai Jiao Tong University, Shanghai, China
Chao Ma
Chao Ma
Professor, Shanghai Jiao Tong University
Computer visionMachine learningImage processing