Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement

📅 2025-04-06
🏛️ IEEE International Conference on Acoustics, Speech, and Signal Processing
📈 Citations: 3
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of conventional personalized speech enhancement methods, which rely on pre-extracted static speaker embeddings that struggle to adapt to target speaker variations during inference and require computationally expensive upstream models for high-quality embeddings. To overcome these challenges, the authors propose a lightweight speaker encoder with only 150K parameters, coupled with a contrastive knowledge distillation strategy tailored for embedding optimization. This approach enables dynamic refinement of speaker representations at inference time by efficiently distilling discriminative features from a complex teacher model. The proposed method achieves significant performance gains in speech enhancement while maintaining minimal computational overhead.

Technology Category

Application Category

📝 Abstract
Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorporate a representation of the target voice within the enhancement system, which is extracted from an enrollment clip of the target voice with upstream models. Those models are generally heavy as the speaker embedding’s quality directly affects PSE performances. Yet, embeddings generated beforehand cannot account for the variations of the target voice during inference time. In this paper, we propose to perform on-the-fly refinement of the speaker embedding using a tiny speaker encoder. We first introduce a novel contrastive knowledge distillation methodology in order to train a 150k-parameter encoder from complex embeddings. We then use this encoder within the enhancement system during inference and show that the proposed method greatly improves PSE performances while maintaining a low computational load.
Problem

Research questions and friction points this paper is trying to address.

Personalized Speech Enhancement
Speaker Embedding
Embedding Refinement
Voice Variability
Inference-time Adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contrastive Knowledge Distillation
On-the-fly Embedding Refinement
Personalized Speech Enhancement
Lightweight Speaker Encoder
Speaker Embedding Optimization
🔎 Similar Papers
No similar papers found.
T
Thomas Serre
Signal Processing and Machine Learning department, Orosound, Paris, France
Mathieu Fontaine
Mathieu Fontaine
Associate Professor, Télécom ParisTech
Machine ListeningSpeech ProcessingDeep Neural Network
É
Éric Benhaim
Signal Processing and Machine Learning department, Orosound, Paris, France
S
S. Essid
LTCI, Télécom Paris, Institut polytechnique de Paris, Palaiseau, France