PhasorNet: Learning Structure from Frequency for Real-Time Stereo Matching

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决立体匹配中细结构等区域的挑战,提出PhasorNet框架,利用频率域线索增强几何区分,并通过多尺度边缘引导高误差区域损失优化训练。
📝 Abstract
Accurate stereo matching remains challenging in ill-posed regions such as fine structures, reflective, or transparent objects, where appearance cues are often ambiguous or unreliable. To tackle this, we propose PhasorNet, a lightweight yet powerful framework that boosts geometric discrimination via frequency-domain cues. At its core, the Phase-Augmented Transformer (PAT) injects Fourier-derived phase information into the attention mechanism, yielding photometrically robust, structure-preserving features that prioritize structural consistency in difficult areas. Additionally, we develop a Geometry-Context Fusion Refinement Module (GCFRM) that combines a full-resolution convolutional stream with a lightweight attention-based stream (leveraging WQA and CDGA blocks) to efficiently preserve fine details and object boundaries without excessive overhead. Training is further enhanced by a multi-scale Edge-guided High-Error Region (EHR) loss that adaptively focuses optimization on high-error and edge regions, guiding hierarchical cost volume refinement. With only 5.3M parameters, PhasorNet achieves state-of-the-art performance on the challenging ETH3D benchmark while exhibiting excellent cross-domain generalization on KITTI, delivering an efficient and practical solution for accurate real-time stereo matching.
Problem

Research questions and friction points this paper is trying to address.

stereo matching
ill-posed regions
fine structures
reflective objects
transparent objects
Innovation

Methods, ideas, or system contributions that make the work stand out.

Phase-Augmented Transformer
Geometry-Context Fusion Refinement Module
Edge-guided High-Error Region loss
🔎 Similar Papers
No similar papers found.