R-LiViT: A LiDAR-Visual-Thermal Dataset Enabling Vulnerable Road User Focused Roadside Perception
Detecting vulnerable road users (VRUs) in roadside perception remains challenging under extreme lighting conditions due to occlusions and visual degradation. Method: We introduce the first roadside multimodal dataset specifically designed for VRU perception, concurrently capturing LiDAR, RGB, and thermal infrared data across three intersections under both daytime and nighttime conditions—encompassing over 150 real-world scenarios. We propose a novel pipeline for precise spatiotemporal alignment of all three modalities, including timestamp synchronization, spatial calibration, and frame-level cross-modal registration. Annotations include fine-grained VRU instance labels: six categories for LiDAR point clouds and eight for image modalities. Contribution/Results: The dataset comprises 10,000 annotated LiDAR frames and 2,400 aligned RGB–thermal image pairs, supporting detection and tracking tasks. All data and evaluation code are publicly released, addressing the absence of thermal imaging in roadside VRU perception and establishing a new benchmark for multimodal roadside intelligent sensing.