Cross-View Sequential Visual Localization with Spatio-Temporal Context Modeling for Autonomous Driving

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing cross-view visual localization methods, which often neglect temporal information and consequently suffer degraded performance under dynamic occlusions, illumination variations, and repetitive textures. To overcome this, we propose the first sequential localization framework that integrates spatiotemporal context by recursively aggregating historical states through a cross-frame module to enhance ground-level features in the current frame. Coupled with hierarchical fine-grained feature representations, our approach enables robust candidate region classification and precise offset estimation from satellite imagery. By introducing temporal modeling into cross-view localization for the first time, our method significantly improves both robustness and accuracy over single-frame baselines. Experiments demonstrate a reduction in average localization error from 3.80 m to 1.57 m on the CVIS dataset, with R@1m improving from 8.14% to 40.22%. Moreover, it achieves zero-shot transfer to KITTI-CVL with a 2.61 m error and yields a real-world vehicle test error of 2.84 m, attaining R@5m of 96.86%.
📝 Abstract
Continuous and reliable localization is essential for autonomous driving. Cross-view visual localization matches ground images with satellite maps, providing complementary localization cues for pipelines that depend on Global Navigation Satellite System (GNSS) signals and high-definition (HD) maps. Most existing cross-view visual localization methods process each frame independently, leaving temporal information underused and limiting accuracy under dynamic occlusion, illumination variation, and repetitive textures. This study proposes a temporal-context-enhanced framework for cross-view sequence visual localization. The proposed recurrent cross-frame module aggregates historical context from the previous state to enhance the coarse ground feature of each current frame. These enhanced features facilitate satellite candidate-region classification, while hierarchical fine-grained features enable precise local offset estimation. On the CVIS dataset, the proposed method reduces mean localization error from 3.80 m to 1.57 m and increases R@1 m from 8.14% to 40.22%. Direct transfer to KITTI-CVL achieves a mean error of 2.61 m, with target-domain fine-tuning further reducing the mean error to 2.27 m. Zero-shot field experiments on a real-world vehicle achieve a mean error of 2.84 m and R@5 m of 96.86%. These results demonstrate that temporal context enhancement significantly improves cross-view localization accuracy and supports robust deployment on public benchmarks and real-world roads.
Problem

Research questions and friction points this paper is trying to address.

cross-view localization
temporal context
visual localization
autonomous driving
spatio-temporal modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-view localization
temporal context modeling
recurrent cross-frame module
visual localization
autonomous driving
🔎 Similar Papers
No similar papers found.