🤖 AI Summary
This study addresses the challenge of accurately estimating the relative transfer matrix (ReTM) in multi-source, multi-microphone scenarios, where conventional covariance-based methods suffer from performance limitations that hinder speech enhancement in noisy environments. To overcome this, the work introduces, for the first time, a deep learning approach to ReTM estimation, proposing an end-to-end trainable supervised learning framework. The model employs time-domain convolution, short-time Fourier transform (STFT)-domain convolution, and LSTM networks to directly learn the ReTM from multi-channel recordings. By circumventing the reliance on covariance matrices inherent in traditional methods, the proposed approach achieves significantly better performance across five objective metrics and demonstrates speech enhancement results comparable to state-of-the-art baselines.
📝 Abstract
The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and sources, shows promising performance when applied to speech enhancement in noisy environments. Estimating the ReTM of sound sources by exploiting the covariance matrices of multichannel recordings is highly beneficial for practical applications and, to date, remains the only proposed approach. This paper investigates deep learning-based ReTM estimation. We propose three novel supervised learning frameworks using time and short-time frequency transform domain convolutional networks, and a Long Short-Term Memory-based recurrent neural network. Experimental results demonstrate that the proposed models achieve more accurate estimation of the ReTM using five objective metrics compared to the covariance-based method. We also show the effectiveness of the proposed frameworks for speech enhancement, achieving performance on par with the baseline method.