Predictive audio representations for early detection and tracking of hidden dynamic objects

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过自监督预训练和多任务微调的方法,利用音频信号预测隐藏动态物体的数量、类型及方向,以解决自动驾驶中被遮挡交通参与者难以及时检测的问题。
📝 Abstract
Predicting potential dangers is core to safety. Forecasting the presence of other traffic agents is core to danger prediction. Occluded traffic agents challenge detection systems as they might become visible too late, leaving the autonomous vehicle too little time to identify, plan and act accordingly in a robust and safe way. Previous works proved that auditory perception, being omnidirectional and not constrained by a field-of-view, provides fundamental cues for early spotting of different road users, even when hidden by other vehicles or infrastructures. Yet, those contributions deal with scenarios with only one vehicle present, and they either identify the type of vehicle or estimate its direction of arrival. In this work, we move forward, and propose a multi-task system which, simultaneously, estimates the number of vehicles present, their type, and their direction of arrival. Our methodological contribution is a two-stage pipeline: a self-supervised pre-training stage inspired by the Joint- Embedding Predictive Architecture (JEPA) applied directly to multichannel raw waveforms, followed by supervised multi-task fine-tuning with a bidirectional LSTM and three classification heads. The pre-training stage trains the encoder without labels, pushing it to predict the latent representation of a future audio segment from its past context. To train and test our framework, since no suitable dataset was publicly available, we collected an ad-hoc one covering Non-Line-Of-Sight scenarios with multiple traffic agents simultaneously operating. Experimental evaluation shows that our method outperforms the state of the art; it also proves that our design choice allows the model to learn robust representations, which can be transferred to an unseen driving scenario, maintaining reasonable and stable performance.
Problem

Research questions and friction points this paper is trying to address.

Occluded traffic agents
autonomous vehicle
audio perception
early detection
multi-task system
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-task System
Self-supervised Pre-training
JEPA
Bidirectional LSTM
🔎 Similar Papers
2024-07-18IEEE Workshop/Winter Conference on Applications of Computer VisionCitations: 0
K
Katerina Vinciguerra
Department of Engineering and Architecture, University of Parma, Parma, Italy
M
Moritz Brandes
Fraunhofer Institute for Digital Media Technology IDMT, Oldenburg, Germany
D
Danilo Hollosi
Fraunhofer Institute for Digital Media Technology IDMT, Oldenburg, Germany
L
Letizia Marchegiani
Department of Engineering and Architecture, University of Parma, Parma, Italy