Institution profile

Cerence

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

A Steered Response Power Method for Sound Source Localization With Generic Acoustic Models

Sep 19, 2025

Traditional SRP methods rely on idealized assumptions—including far-field propagation, omnidirectional sources, and spatially uncorrelated noise—leading to significant degradation in localization accuracy under realistic acoustic conditions. To address this, we propose a generalized SRP beamforming framework grounded in a comprehensive acoustic model. Our approach is the first to explicitly incorporate measured acoustic transfer functions, source/microphone directivity patterns, and acoustic shadowing effects into the SRP formulation, thereby relaxing the restrictive far-field and omnidirectional assumptions. We further design a generalized cost function tailored for spatially correlated noise, jointly exploiting both time-difference-of-arrival (TDOA) and level-difference-of-arrival (LDOA) cues. Additionally, we introduce delay-and-sum optimization and frequency-domain weighting to enhance robustness. Experiments across diverse microphone array geometries and high-noise environments demonstrate superior performance: the proposed method achieves over 60% reduction in average localization error compared to conventional SRP.

0 citationsRead paper

Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation

May 25, 2025

Existing self-supervised speech representations (e.g., WavLM) struggle to fully disentangle speaker identity from linguistic content, thereby degrading downstream task performance. To address this, we propose a lightweight, interpretable linear decomposition framework: leveraging learnable projections with orthogonality constraints, WavLM features are explicitly decomposed into speaker-dependent and speaker-independent components; these are jointly optimized via speaker discrimination loss and content reconstruction objective. Crucially, our method requires no architectural complexity or auxiliary annotations, and—uniquely—achieves *exact*, *lossless*, and *purely linear* speaker disentanglement. Evaluated on voice conversion, it substantially surpasses state-of-the-art methods: speaker similarity decreases by 62%, while speech quality (MOS) and content accuracy (WER) both improve significantly. Inference overhead is negligible.

0 citationsRead paper
Recent publications

Latest Papers

A Steered Response Power Method for Sound Source Localization With Generic Acoustic Models

Sep 19, 2025

Traditional SRP methods rely on idealized assumptions—including far-field propagation, omnidirectional sources, and spatially uncorrelated noise—leading to significant degradation in localization accuracy under realistic acoustic conditions. To address this, we propose a generalized SRP beamforming framework grounded in a comprehensive acoustic model. Our approach is the first to explicitly incorporate measured acoustic transfer functions, source/microphone directivity patterns, and acoustic shadowing effects into the SRP formulation, thereby relaxing the restrictive far-field and omnidirectional assumptions. We further design a generalized cost function tailored for spatially correlated noise, jointly exploiting both time-difference-of-arrival (TDOA) and level-difference-of-arrival (LDOA) cues. Additionally, we introduce delay-and-sum optimization and frequency-domain weighting to enhance robustness. Experiments across diverse microphone array geometries and high-noise environments demonstrate superior performance: the proposed method achieves over 60% reduction in average localization error compared to conventional SRP.

0 citationsRead paper

Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation

May 25, 2025

Existing self-supervised speech representations (e.g., WavLM) struggle to fully disentangle speaker identity from linguistic content, thereby degrading downstream task performance. To address this, we propose a lightweight, interpretable linear decomposition framework: leveraging learnable projections with orthogonality constraints, WavLM features are explicitly decomposed into speaker-dependent and speaker-independent components; these are jointly optimized via speaker discrimination loss and content reconstruction objective. Crucially, our method requires no architectural complexity or auxiliary annotations, and—uniquely—achieves *exact*, *lossless*, and *purely linear* speaker disentanglement. Evaluated on voice conversion, it substantially surpasses state-of-the-art methods: speaker similarity decreases by 62%, while speech quality (MOS) and content accuracy (WER) both improve significantly. Inference overhead is negligible.

0 citationsRead paper