When Speech Meets Lips: Interpretable Audio-Visual Synchronization for L2 Pronunciation Assessment

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究提出一种音频-视觉同步框架,通过特征编码、交叉注意力融合等方法解决二语发音评估中语音与唇部动作时间同步问题。
📝 Abstract
Automatic Pronunciation Assessment (APA) systems have achieved strong performance with transformer-based models and self-supervised speech representations. However, most methods rely only on acoustic signals and overlook temporal synchronization between speech and articulatory movements, limiting diagnostic feedback on timing mismatches important for L2 pronunciation training. We propose an interpretable audio-visual synchronization framework that explicitly models speech-lip temporal alignment through feature encoding, cross-attention fusion, lag estimation, stability quantification, and visualization. The framework introduces frame-level lag trajectories and a Lag Stability Index (LSI) to quantify synchronization robustness. We also interviewed 30 participants, including 10 instructors and 20 students with diverse first-language backgrounds, to assess its effectiveness. By transforming implicit alignment into interpretable representations, the framework connects automatic scoring with actionable Computer-Aided Pronunciation Training feedback. Datasets and supplemental materials are available at https://www.robots.ox.ac.uk/~vgg/data/lip_reading/.
Problem

Research questions and friction points this paper is trying to address.

Automatic Pronunciation Assessment
audio-visual synchronization
speech-lip temporal alignment
L2 pronunciation training
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio-visual synchronization
lag estimation
Lag Stability Index (LSI)
cross-attention fusion
interpretable representations
💼 Related Jobs
No related jobs found.