Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该论文提出使用神经控制微分方程解决TTS中的时长感知声学建模问题,通过连续时间机制生成随音素内容和时长变化的隐藏状态。
📝 Abstract
Text-to-speech (TTS) models commonly address text--speech alignment by expanding phone-level encoder states to frame-level decoder inputs using predicted durations. While this length-regulation step resolves alignment structurally, this use of duration typically changes only where and how often latent states appear, not the values of the states themselves. This paper proposes a continuous-time mechanism for duration-aware acoustic modelling in TTS using neural controlled differential equations (CDEs). We formulate the phone representation as a temporally parameterised control path and use a neural acoustic vector field to produce a continuous-time hidden state whose values evolve with phonetic content and duration-derived timing. The resulting trajectory can be sampled at discrete points and integrated into a standard acoustic decoder pipeline. Objective results contrast CDEs and typical recurrent models. Subjective results suggest that CDE-based models evaluating one phone per step can improve rank-order agreement between synthesised and reference emotion intensity while maintaining comparable emotion-expression quality to a strong baseline. Additional experiments with half-phone step-sizes suggest that temporal resolution changes the trade-off between style tracking and absolute calibration. These results position CDEs as a promising design space for continuous-time and duration-aware style-sensitive TTS.
Problem

Research questions and friction points this paper is trying to address.

Text-to-speech
duration-aware
acoustic modelling
neural controlled differential equations
temporal resolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

neural controlled differential equations
continuous-time acoustic modelling
duration-aware TTS
temporal resolution
emotion intensity
🔎 Similar Papers
No similar papers found.
M
Mattias Cross
Speech and Hearing Group, University of Sheffield, Sheffield, United Kingdom
M
Minghui Zhao
Speech and Hearing Group, University of Sheffield, Sheffield, United Kingdom
Anton Ragni
Anton Ragni
University of Sheffield
Speech and Language Technologies