🤖 AI Summary
To address unstable feature detection and tracking in large-scale, long-term outdoor visual odometry (VO) caused by illumination variations, dynamic scenes, and low-texture regions, this paper proposes a task-driven self-supervised feature learning framework. Unlike existing supervised approaches relying on SuperPoint/SuperGlue, our method employs VO motion estimation error as a feedback signal to iteratively optimize feature detection, description, and matching in an end-to-end self-supervised training paradigm. This closed-loop optimization significantly improves feature robustness and out-of-distribution generalization. Experimental results demonstrate that the proposed method enhances feature tracking stability by 23.6% and improves VO localization accuracy by 18.4% on challenging real-world sequences—particularly excelling in low-texture and strongly varying illumination conditions.
📝 Abstract
Visual-based localization has made significant progress, yet its performance often drops in large-scale, outdoor, and long-term settings due to factors like lighting changes, dynamic scenes, and low-texture areas. These challenges degrade feature extraction and tracking, which are critical for accurate motion estimation. While learning-based methods such as SuperPoint and SuperGlue show improved feature coverage and robustness, they still face generalization issues with out-of-distribution data. We address this by enhancing deep feature extraction and tracking through self-supervised learning with task specific feedback. Our method promotes stable and informative features, improving generalization and reliability in challenging environments.