Efficient Sequential Neural Network with Spatial-Temporal Attention and Linear LSTM for Robust Lane Detection Using Multi-Frame Images
This work addresses the challenges of achieving accuracy, robustness, and real-time performance in visual lane detection under complex scenarios such as occlusion and strong illumination, where existing methods often fail to exploit the spatiotemporal saliency of critical regions. To this end, we propose a lightweight encoder-decoder sequential network that integrates a spatial-temporal attention mechanism with a linear LSTM module. By leveraging multi-frame inputs, our model effectively captures spatiotemporal dependencies and adaptively focuses on salient lane features. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art approaches on three large-scale public benchmarks, particularly excelling in challenging conditions, while significantly reducing both model parameters and computational cost (measured in MACs), thereby enabling efficient and robust end-to-end multi-frame lane detection.