MLG-Stereo: ViT Based Stereo Matching with Multi-Stage Local-Global Enhancement

📅 2026-04-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing Vision Transformer (ViT)-based stereo matching methods, which suffer from resolution sensitivity and inadequate modeling of local details, hindering their ability to efficiently handle images of arbitrary resolutions. To overcome these challenges, the authors propose MLG-Stereo, a novel framework that integrates a local-global enhancement mechanism throughout the entire pipeline—spanning feature extraction, cost volume construction, and disparity refinement—for the first time. Built upon a ViT architecture, MLG-Stereo incorporates a multi-granularity feature network, a local-global cost volume, and a local-global guided recurrent unit to sustain global modeling capacity beyond the encoder. Experiments demonstrate that MLG-Stereo achieves state-of-the-art performance on KITTI 2012 and highly competitive results on Middlebury and KITTI 2015, effectively bridging the scale gap between training and inference.

Technology Category

Application Category

📝 Abstract
With the development of deep learning, ViT-based stereo matching methods have made significant progress due to their remarkable robustness and zero-shot ability. However, due to the limitations of ViTs in handling resolution sensitivity and their relative neglect of local information, the ability of ViT-based methods to predict details and handle arbitrary-resolution images is still weaker than that of CNN-based methods. To address these shortcomings, we propose MLG-Stereo, a systematic pipeline-level design that extends global modeling beyond the encoder stage. First, we propose a Multi-Granularity Feature Network to effectively balance global context and local geometric information, enabling comprehensive feature extraction from images of arbitrary resolution and bridging the gap between training and inference scales. Then, a Local-Global Cost Volume is constructed to capture both locally-correlated and global-aware matching information. Finally, a Local-Global Guided Recurrent Unit is introduced to iteratively optimize the disparity locally under the guidance of global information. Extensive experiments are conducted on multiple benchmark datasets, demonstrating that our MLG-Stereo exhibits highly competitive performance on the Middlebury and KITTI-2015 benchmarks compared to contemporaneous leading methods, and achieves outstanding results in the KITTI-2012 dataset.
Problem

Research questions and friction points this paper is trying to address.

stereo matching
Vision Transformer
local-global information
resolution sensitivity
detail prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

ViT-based stereo matching
multi-granularity feature
local-global cost volume
recurrent disparity refinement
arbitrary-resolution generalization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Haoyu Zhang
Haoyu Zhang
Ph.D. candidate, Norwegian University of Science and Technology
J
Jingyi Zhou
Embedded Deep Learning and Visual Analysis Laboratory, College of Future Information Technology, Fudan University, Shanghai 200433, China
Peng Ye
Peng Ye
LIDYL, CEA, University Paris-Saclay
Attosecond ScienceStrong FieldUltrafast OpticsHHG in gas and solid
Jiakang Yuan
Jiakang Yuan
Fudan university
MLLMsMulti-agent SystemReasoning
Lin Zhang
Lin Zhang
Fudan University
F
Feng Xu
Key Laboratory for Information Science of Electromagnetic Waves, Ministry of Education, Fudan University, Shanghai 200433, China; College of Future Information Technology, Fudan University, Shanghai 200433, China
Tao Chen
Tao Chen
Fudan University
Deep LearningMedical Image Segmentation