Optimizing Multilingual Text-To-Speech with Accents&Emotions

📅 2025-06-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address accent distortion, emotional incoherence, and insufficient cultural adaptation in text-to-speech (TTS) systems for Hindi and Indian English synthesis within South Asia’s multilingual context, this work proposes the first accent–emotion disentangled architecture. Our method integrates a language-specific phoneme-aligned hybrid encoder-decoder with residual vector-quantized accent encoding, enabling real-time cross-lingual accent switching (e.g., “Namaste, let’s talk about”) and culture-aware emotional embedding. Building upon Parler-TTS, we train a culturally sensitive emotion layer and a dynamic accent code switching module on native speech corpora. Experiments demonstrate a 23.7% improvement in accent accuracy (WER reduced from 15.4% to 11.8%), 85.3% native speaker emotion recognition accuracy, and a cultural correctness MOS of 4.2/5 (p < 0.01), significantly outperforming METTS and VECL-TTS.

Technology Category

Application Category

📝 Abstract
State-of-the-art text-to-speech (TTS) systems realize high naturalness in monolingual environments, synthesizing speech with correct multilingual accents (especially for Indic languages) and context-relevant emotions still poses difficulty owing to cultural nuance discrepancies in current frameworks. This paper introduces a new TTS architecture integrating accent along with preserving transliteration with multi-scale emotion modelling, in particularly tuned for Hindi and Indian English accent. Our approach extends the Parler-TTS model by integrating A language-specific phoneme alignment hybrid encoder-decoder architecture, and culture-sensitive emotion embedding layers trained on native speaker corpora, as well as incorporating a dynamic accent code switching with residual vector quantization. Quantitative tests demonstrate 23.7% improvement in accent accuracy (Word Error Rate reduction from 15.4% to 11.8%) and 85.3% emotion recognition accuracy from native listeners, surpassing METTS and VECL-TTS baselines. The novelty of the system is that it can mix code in real time - generating statements such as"Namaste, let's talk about"with uninterrupted accent shifts while preserving emotional consistency. Subjective evaluation with 200 users reported a mean opinion score (MOS) of 4.2/5 for cultural correctness, much better than existing multilingual systems (p<0.01). This research makes cross-lingual synthesis more feasible by showcasing scalable accent-emotion disentanglement, with direct application in South Asian EdTech and accessibility software.
Problem

Research questions and friction points this paper is trying to address.

Improving multilingual TTS accent accuracy for Indic languages
Enhancing emotion synthesis with cultural nuance sensitivity
Enabling real-time code-switching with preserved emotional consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid encoder-decoder for phoneme alignment
Culture-sensitive emotion embedding layers
Dynamic accent code switching with quantization
P
Pranav Pawar
Dwarkadas J. Sanghvi College of Engineering, Mumbai, India
A
Akshansh Dwivedi
Dwarkadas J. Sanghvi College of Engineering, Mumbai, India
J
Jenish Boricha
Dwarkadas J. Sanghvi College of Engineering, Mumbai, India
H
Himanshu Gohil
Dwarkadas J. Sanghvi College of Engineering, Mumbai, India
A
Aditya Dubey
Dwarkadas J. Sanghvi College of Engineering, Mumbai, India