Multimodal Emergency Vehicle Classification via Audio-Visual Transformers and Knowledge Distillation

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决单一模态在复杂环境下失效的问题,本文提出AVNet,利用音频-视觉变换器并通过知识蒸馏方法提高紧急车辆分类的准确性。
📝 Abstract
Emergency vehicle detection in autonomous driving is a safety-critical perception task that demands robustness under diverse and adverse real-world conditions. Existing approaches rely on a single modality, either audio or video, which leads to systematic failure when that modality is degraded: microphone-based systems fail in noisy urban environments, and camera-based systems fail at night or under occlusion. This report presents AVNet, a multimodal audio-visual transformer that classifies emergency vehicles (ambulance, fire engine, police car) and road background using both audio and video, while gracefully handling the absence of either modality at inference time. AVNet introduces three key contributions: (1) a temporally aligned cross-modal fusion module that performs second-level cross-attention between audio spectrogram tokens and video frame tokens, exploiting their exact temporal correspondence without any learned alignment mechanism; (2) learned null embeddings that substitute for missing modality tokens, enabling a single unified model to operate in audio-only, video-only, or joint audio-visual mode without retraining; and (3) a knowledge distillation training strategy in which specialist unimodal teacher models transfer inter-class dark knowledge into the multimodal student fusion branch via soft probability targets. Evaluated on 281 clips from the Google AudioSet dataset, AVNet achieves 66.6% overall accuracy in audio-visual mode, outperforming the audio-only branch by +10.4% and the video-only branch by +15.0%. The largest per-class gain is observed for the hardest class, Ambulance, where fusion achieves +29.5% over either unimodal branch alone, demonstrating that the two modalities provide complementary information that the aligned cross attention mechanism successfully exploits.
Problem

Research questions and friction points this paper is trying to address.

Emergency Vehicle Detection
Autonomous Driving
Multimodal Perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Transformers
Cross-modal Fusion
Knowledge Distillation
Null Embeddings
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Vijay John
Vijay John
Lawrence Technological University
Intelligent MobilityGuardian RoboticsMultimodal Sensor Fusion and CalibrationHuman Motion
A
Amar Dabaja
Department of Mathematics and Computer Science, Lawrence Technological University, Southfield, MI, USA