Training Set Synthesis for Bioacoustic Denoising: A Case Study With Mice

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of denoising bioacoustic recordings corrupted by environmental noise, where clean reference signals are typically unavailable for supervised training. To circumvent the need for real clean data, the authors propose a self-supervised approach that synthesizes training samples containing fundamental frequency and harmonic ridges. They develop a U-Net-based model to predict complex ratio masks and introduce a ridge-guided weighted loss function to better preserve fine-grained vocal structure during denoising. Evaluated on murine ultrasonic vocalizations, the method significantly improves the accuracy of fundamental frequency and harmonic tracking, enhances scale-invariant signal-to-noise ratio, and boosts the generalization performance of downstream classifiers in noisy field conditions.
📝 Abstract
Bioacoustic recordings are often degraded by ambient noise, which complicates the analysis of weak or noise-overlapped vocalizations. Convolutional neural networks, particularly U-Net architectures, have shown a strong denoising performance in speech and music processing. However, their direct application to bioacoustic signals is limited by the scarcity of clean training data. To address this issue, we propose a training set synthesis approach and develop a supervised denoising model that predicts a complex ratio mask in the time-frequency domain. The model leverages ridges, or frequency contours, that represent the fundamental frequency together with one or more harmonic partial components of vocalizations. These ridges are used both for the synthesis of training sets and to design a loss function that assigns higher weights to the ridge regions (ridge-guided loss function). This weighting step helps the network better preserve vocalization details during denoising. As a case study, we evaluate our approach using ultrasonic vocalizations (USVs) recordings of house mice, which are widely studied in behavioral biology and neuroscience. In actual field recordings, the proposed method enhances fundamental and harmonic partial ridge tracking compared to our previous signal-processing approach. In addition, a classifier trained on denoised data improves USV classification on out-of-sample, noisy recordings from wild and domesticated mice compared to classifiers trained on noisy recordings. Our proposed method also substantially improves the scale-invariant signal-to-distortion ratio on synthetic testing data across a wide range of input signal-to-noise ratios. Although we focus on USVs, the proposed approach should be broadly applicable to other bioacoustic signals with trackable ridges, and thus enables ridgebased training set synthesis and denoising.
Problem

Research questions and friction points this paper is trying to address.

bioacoustic denoising
training data scarcity
ultrasonic vocalizations
ambient noise
vocalization analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

training set synthesis
ridge-guided loss
bioacoustic denoising
complex ratio mask
ultrasonic vocalizations
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
R
Reyhaneh Abbasi
Acoustics Research Institute of the Austrian Academy of Sciences, Vienna, Austria
Peter Balazs
Peter Balazs
Acoustics Research Institute, Austrian Academy of Sciences
Application-oriented mathematicsAcousticsFrame Theory
Vincent Lostanlen
Vincent Lostanlen
LS2N, CNRS
Machine listening
C
Clara Hollomey
Institute for Creative Media Technologies, University of Applied Sciences Saint Pölten
D
Dustin J. Penn
Konrad Lorenz Institute of Ethology, Department of Interdisciplinary Life Sciences, University of Veterinary Medicine, Vienna, Austria
S
Sarah M. Zala
Konrad Lorenz Institute of Ethology, Department of Interdisciplinary Life Sciences, University of Veterinary Medicine, Vienna, Austria
Nicki Holighaus
Nicki Holighaus
Acoustics Research Institute, Austrian Academy of Sciences