LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of dense matching in global multi-temporal remote sensing imagery, which suffers from large geometric offsets, partial overlaps, and non-matchable regions due to varying imaging conditions. The authors propose a two-stage “localization–registration” framework: first, a matchability-aware module identifies matchable regions and estimates an affine transformation; subsequently, dense residual correspondences are predicted within the aligned coordinate system under guided supervision. Key contributions include the creation of LEVIR-GM—the first global remote sensing matching benchmark with native matchability annotations—and a unified evaluation protocol, along with an efficient joint architecture that fuses multi-scale features. On LEVIR-GM, the method achieves an AUC of 83.3%, outperforming RoMa v2 by 1.6 points, improves PCK at ½-pixel by 6.5 and 8.2 points, reduces inference latency by 47.8%, and demonstrates strong generalization in cross-platform geolocation tasks.
📝 Abstract
Dense image matching establishes pixel-wise correspondences and underpins broad applications in computer vision and photogrammetry. However, extending dense matching to global-scale remote sensing remains challenging because image pairs may differ in acquisition time, season, viewpoint, spatial resolution, and land-cover state. The resulting large geometric offsets, partial overlap, and intrinsically unmatchable regions make direct dense correspondence prediction unreliable and inefficient. We thus reformulate dense matching as localization-and-registration: first localizing the matchable overlap and affine geometry, then refining dense residuals within the aligned frame. Based on this formulation, we propose LoRetta, a foundation model coupling matchability-aware affine localization with guided dense registration. We also introduce LEVIR-GM, a global-scale multi-temporal optical matching benchmark with dataset-native matchability labels (103K aligned, 827K augmented pairs, six continents, five years, 0.5-1024 m resolution). We further establish a unified evaluation protocol for sparse, semi-dense, and dense matchers. On LEVIR-GM, LoRetta achieves an area under the curve (AUC) of 83.3%, outperforming the strongest baseline RoMa v2 by 1.6 points, with larger percentage of correct keypoints (PCK) gains of 6.5 and 8.2 points at 1 and 2 pixels, while reducing inference latency by 47.8%. Astronaut-to-satellite and unmanned aerial vehicle (UAV)-to-satellite geolocalization experiments further demonstrate its transferability as a reusable geometric aligner.
Problem

Research questions and friction points this paper is trying to address.

dense image matching
remote sensing
global-scale
geometric correspondence
multi-temporal
Innovation

Methods, ideas, or system contributions that make the work stand out.

dense image matching
foundation model
affine localization
remote sensing
geometric registration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Siwei Yu
Department of Aerospace Intelligent Science and Technology, School of Astronautics, Beihang University, Beijing 100191, China, and also with the Key Laboratory of Spacecraft Design Optimization and Dynamic Simulation Technologies, Ministry of Education
Han Guo
Han Guo
Beihang University
Remote SensingComputer Vision
Zhenwei Shi
Zhenwei Shi
Professor at Image Processing Center, Beihang University, China
Hyperspectral imagingRemote SensingSignal and Image ProcessingPattern RecognitionMachine Learning
Zhengxia Zou
Zhengxia Zou
Beihang Univeristy
computer visionimage processingremote sensinggames