Training Set Synthesis for Bioacoustic Denoising: A Case Study With Mice
This study addresses the challenge of denoising bioacoustic recordings corrupted by environmental noise, where clean reference signals are typically unavailable for supervised training. To circumvent the need for real clean data, the authors propose a self-supervised approach that synthesizes training samples containing fundamental frequency and harmonic ridges. They develop a U-Net-based model to predict complex ratio masks and introduce a ridge-guided weighted loss function to better preserve fine-grained vocal structure during denoising. Evaluated on murine ultrasonic vocalizations, the method significantly improves the accuracy of fundamental frequency and harmonic tracking, enhances scale-invariant signal-to-noise ratio, and boosts the generalization performance of downstream classifiers in noisy field conditions.