Subphonetic Acoustic Modeling via Optimal Transport for Pronunciation Assessment

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决发音评估中声学模型难以同时提供识别和分割证据的问题,提出结合有序子音素状态与最优时间传输分类的拓扑感知帧级声学模型。
📝 Abstract
Pronunciation assessment requires acoustic evidence that is temporally precise, diagnostically meaningful, and faithful to the learner's actual production. However, existing acoustic models often struggle to provide recognition and segmentation evidence simultaneously. CTC-based phone recognizers can predict phone sequences flexibly, but their sparse and peaky posteriors often miss phone boundaries and fine-grained pronunciation cues. In contrast, text-dependent forced aligners provide reliable temporal information when transcripts are available, but are not directly applicable to reference-free pronunciation analysis. In this work, we propose a topology-aware frame-wise acoustic model that learns dense ordered state posteriors within each phone. The key idea is to recover phone-internal state structure in a neural acoustic model by combining ordered subphonetic states with optimal temporal transport classification (OTTC). This combination encourages dense monotonic frame-level state discrimination while preserving phone recognition ability. Experiments on read, spontaneous, and L2 speech show improved segmentation over neural baselines with competitive recognition performance. Downstream evaluations further show gains in mispronunciation detection and automatic pronunciation assessment. Probing analysis suggests that the learned states capture phoneme-dependent acoustic structure rather than arbitrary frame-level distributions.
Problem

Research questions and friction points this paper is trying to address.

Pronunciation Assessment
Acoustic Modeling
Optimal Transport
Phone Recognition
Temporal Segmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

topology-aware frame-wise acoustic model
ordered subphonetic states
optimal temporal transport classification (OTTC)
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haopeng Geng
Graduate School of Engineering, The University of Tokyo, Japan
J
Jiun-Ting Li
Advanced Technology Laboratory, Chunghwa Telecom Co., Ltd., Taiwan
D
Daisuke Saito
Graduate School of Engineering, The University of Tokyo, Japan
Nobuaki Minematsu
Nobuaki Minematsu
The University of Tokyo
Speech CommunicationForeign Language Learning