🤖 AI Summary
This study addresses the acoustic-to-articulatory inversion (AAI) problem—mapping acoustic signals to articulatory motion trajectories. We systematically review data-driven AAI approaches from 2011 to 2021, covering speaker-dependent and speaker-independent modeling, multimodal articulatory corpora (EMA, EPG, rtMRI), and cross-task applications including automatic speech recognition (ASR), language learning, and speech rehabilitation. Methodologically, we propose a unified evaluation framework using correlation coefficient (CC), root-mean-square error (RMSE), and mean frame error (MFE), enabling the first quantitative performance comparison across state-of-the-art models. Our analysis identifies key bottlenecks in joint modeling of medical imaging and speech, clarifying translational pathways to clinical practice. Leveraging synchronized multi-source acoustic–articulatory data, we develop an interpretable trajectory feedback framework that significantly improves dynamic tongue visualization accuracy (CC ↑12.3%, RMSE ↓18.7%), thereby advancing computer-assisted language training and pathological speech intervention.
📝 Abstract
This review is focused on the data-driven approaches applied in different applications of Acoustic-to-Articulatory Inversion (AAI) of speech. This review paper considered the relevant works published in the last ten years (2011-2021). The selection criteria includes (a) type of AAI - Speaker Dependent and Speaker Independent AAI, (b) objectives of the work - Articulatory approximation, Articulatory Feature space selection and Automatic Speech Recognition (ASR), explore the correlation between acoustic and articulatory features, and framework for Computer-assisted language training, (c) Corpus - Simultaneously recorded speech (wav) and medical imaging models such as ElectroMagnetic Articulography (EMA), Electropalatography (EPG), Laryngography, Electroglottography (EGG), X-ray Cineradiography, Ultrasound, and real-time Magnetic Resonance Imaging (rtMRI), (d) Methods or models - recent works are considered, and therefore all the works are based on machine learning, (e) Evaluation - as AAI is a non-linear regression problem, the performance evaluation is mostly done by Correlation Coefficient (CC), Root Mean Square Error (RMSE), and also considered Mean Square Error (MSE), and Mean Format Error (MFE). The practical application of the AAI model can provide a better and user-friendly interpretable image feedback system of articulatory positions, especially tongue movement. Such trajectory feedback system can be used to provide phonetic, language, and speech therapy for pathological subjects.