Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

📅 2026-08-10
📈 Citations: 0
✹ Influential: 0
📄 PDF
đŸ€– AI Summary
Mapping real-time magnetic resonance imaging (rtMRI)–derived articulatory movements to phonological categories remains challenging. This work proposes a multimodal modeling approach that, for the first time, transfers audio representations trained with phonological feature supervision—based on PhonoQ, WavLM-large, and HuBERT-large—into purely articulatory modeling, yielding performance gains even during audio-free inference. By fusing synchronized audio and articulatory contours, the model significantly improves macro-F1 scores for phonological classification tasks (e.g., manner and place of articulation, voicing, vowel height, and backness) under both unseen-speech and unseen-speaker conditions, while also enhancing fine-grained accuracy across 39 phoneme classes. Furthermore, the approach reveals interpretable patterns of articulatory variation.
📝 Abstract
Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based model trained to recognize structured phonological features, provides useful information for audio--articulatory modeling. Specifically, we extract representations from PhonoQ's Conformer module, whose training is shaped by supervision for manner, place, voicing, and vowel features. Using articulatory contours with synchronized audio-derived features, we compare WavLM-large and HuBERT-large baselines with models that incorporate PhonoQ-derived representations. Across unseen-speech and unseen-subject settings, these features improve macro-F1 for phonological targets including manner, place, voicing, vowel height, and vowel backness, and also improve fine-grained 39-phoneme classification. In a contour-only inference setting, audio-derived teacher supervision yields modest but consistent gains over contour-only training, indicating that phonological information from synchronized audio can be partially transferred to articulatory models. Finally, posterior analyses show interpretable surface-sensitive patterns consistent with flapping-like /t/ realizations, /t/-/r/ retraction or affrication, and nasal place assimilation.
Problem

Research questions and friction points this paper is trying to address.

rtMRI speech classification
phonological representation
audio-articulatory modeling
phoneme classification
articulatory-to-phonetic mapping
Innovation

Methods, ideas, or system contributions that make the work stand out.

structured phonological representations
audio-articulatory modeling
rtMRI speech classification
PhonoQ
feature transfer
🔎 Similar Papers
No similar papers found.
đŸ’Œ Related Jobs
No related jobs found.
A
Abner Hernandez
Pattern Recognition Lab, Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg (FAU), Erlangen, Germany
T
TomĂĄs Arias Vergara
Pattern Recognition Lab, Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg (FAU), Erlangen, Germany; Department of Electronic Engineering, Universidad de Antioquia, MedellĂ­n, Colombia
D
Daiqi Liu
Pattern Recognition Lab, Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg (FAU), Erlangen, Germany
A
Andreas Maier
Pattern Recognition Lab, Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg (FAU), Erlangen, Germany
Paula Andrea Pérez-Toro
Paula Andrea Pérez-Toro
Friedrich-Alexander-UniversitĂ€t Erlangen-NĂŒrnberg; Universidad de Antioquia
Machine LearningSpeech AnalysisGait AnalysisNatural Language ProcessingDeep Learning