🤖 AI Summary
To address the challenge of constructing high-definition semantic maps in complex intersections—where onboard solutions suffer from occlusions and limited field-of-view—this paper proposes a roadside infrastructure-based multimodal mapping method leveraging Intelligent Roadside Units (IRUs). We design a two-stage camera-LiDAR fusion framework that jointly optimizes modality-specific feature extraction and cross-modal semantic alignment, enabling high-fidelity geometric-textural joint modeling. We introduce RS-seq, the first publicly available sequential dataset specifically designed for roadside HD map generation, thereby establishing the first systematic benchmark for this emerging domain. Extensive evaluation on RS-seq demonstrates that our method achieves a semantic segmentation mIoU 4% higher than image-only baselines and 18% higher than point-cloud-only baselines, significantly outperforming existing approaches. This work establishes a new paradigm for high-precision, perception-driven semantic mapping from roadside infrastructure.
📝 Abstract
High-definition (HD) semantic mapping of complex intersections poses significant challenges for traditional vehicle-based approaches due to occlusions and limited perspectives. This paper introduces a novel camera-LiDAR fusion framework that leverages elevated intelligent roadside units (IRUs). Additionally, we present RS-seq, a comprehensive dataset developed through the systematic enhancement and annotation of the V2X-Seq dataset. RS-seq includes precisely labelled camera imagery and LiDAR point clouds collected from roadside installations, along with vectorized maps for seven intersections annotated with detailed features such as lane dividers, pedestrian crossings, and stop lines. This dataset facilitates the systematic investigation of cross-modal complementarity for HD map generation using IRU data. The proposed fusion framework employs a two-stage process that integrates modality-specific feature extraction and cross-modal semantic integration, capitalizing on camera high-resolution texture and precise geometric data from LiDAR. Quantitative evaluations using the RS-seq dataset demonstrate that our multimodal approach consistently surpasses unimodal methods. Specifically, compared to unimodal baselines evaluated on the RS-seq dataset, the multimodal approach improves the mean Intersection-over-Union (mIoU) for semantic segmentation by 4% over the image-only results and 18% over the point cloud-only results. This study establishes a baseline methodology for IRU-based HD semantic mapping and provides a valuable dataset for future research in infrastructure-assisted autonomous driving systems.