Multi-Modal Traffic Sign Detection with Semantic Attributes for Autonomous Driving

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对自动驾驶中交通标志检测的问题,提出了一种结合摄像头和LiDAR的多模态检测框架,通过几何不变量而非区域特定外观特征来提高检测鲁棒性和跨区域适用性。
📝 Abstract
Reliable traffic sign detection is a prerequisite for the global deployment of autonomous driving systems, where regulatory compliance and road safety depend on perceiving signs correctly across regions, ranges, and weather conditions. Despite recent progress, vision-based methods continue to face three fundamental limitations: poor cross-regional generalization due to high diversity across countries, degraded performance on small-object detection at long ranges (traffic signs occupy as little as $10{\times}10$ pixels at 200m), and fragile temporal tracking under the strongly non-linear perspective distortion that occurs as a vehicle approaches a sign. In this paper, we address the problem of robust, long-range, region-agnostic traffic sign perception by combining camera and Light Detection and Ranging (LiDAR) sensing. We present a multi-modal detection framework whose Intensity-Aware Deformable Fusion module aligns retro-reflective LiDAR cues with camera features, anchoring detection on geometric invariants rather than region-specific visual appearance. We further introduce a dual motion-model tracker that explicitly accounts for non-linear perspective transformations during vehicle approach, substantially improving temporal consistency over linear motion assumptions. Additionally, we develop a semantic attribute classification pipeline that estimates occlusion level, readability, sign embeddedness, and road relevance, providing actionable context to downstream planning. Extensive evaluation on our dataset, spanning 60+ countries and 2,500+ hours of driving data, shows that the proposed pipeline achieves an Object Miss Ratio (OMR) of 0.49% across 221,068 evaluation sequences, demonstrating globally generalizable traffic sign perception in commercial-grade autonomous driving systems.
Problem

Research questions and friction points this paper is trying to address.

traffic sign detection
cross-regional generalization
long-range detection
non-linear perspective distortion
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-modal detection
Intensity-Aware Deformable Fusion
dual motion-model tracker
semantic attribute classification
🔎 Similar Papers
No similar papers found.
M
Meda Lazar
Arriver System Software S.r.l., Romania
S
Sourab Sridhar
Qualcomm Auto Ltd Sweden Filial
S
Shashwata Gupta
Qualcomm Auto Ltd., UK
A
Alexandra Tripcea
Arriver System Software S.r.l., Romania
V
Varun Ravi
Automated Driving, Qualcomm Technologies, Inc
Senthil Yogamani
Senthil Yogamani
Engineering Director, Data-centric AI for Autonomous Driving at Qualcomm Inc
Multimodal PerceptionData-centric AIIntelligent VehiclesAutonomous DrivingVLA