Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of comprehensive evaluation frameworks for cross-lingual text-to-speech (TTS) systems in low-resource languages that jointly assess perceptual quality, speaker similarity, and acoustic fidelity. The authors propose a reproducible multi-metric benchmark integrating MUSHRA/ABX subjective listening tests, Resemblyzer-based speaker similarity scores, and objective measures such as Mel-cepstral distortion (MCD) and F0 RMSE. For the first time, four state-of-the-art TTS systems are rigorously evaluated across four distinct speech domains—formal, conversational, literary, and emotional—in a low-resource setting. Results reveal that emotional speech synthesis poses the greatest challenge (average MCD: 12.03 dB), while conversational speech achieves the highest acoustic fidelity, with significant performance variations observed across systems and domains. The complete evaluation toolkit and dataset are publicly released to advance standardized assessment in low-resource TTS research.
📝 Abstract
Recent advances in neural text-to-speech (TTS) systems have substantially improved speech naturalness and intelligibility across many languages. However, comprehensive evaluation methodologies that jointly assess perceptual quality, speaker similarity, and acoustic fidelity across diverse speech domains remain limited, particularly for low-resource and underrepresented languages. This paper presents a reproducible, multi-metric benchmarking framework for systematic evaluation of modern TTS systems through domain-specific analysis. The proposed framework integrates complementary subjective and objective evaluation protocols and is demonstrated through a comprehensive case study on a representative low-resource language spanning four speech domains: Formal, Conversational, Literary/Storytelling, and Emotional. Four state-of-the-art TTS systems -- Indic-Parler-TTS, MMS-TTS, Microsoft Edge TTS, and Google Gemini TTS -- are evaluated using MUSHRA listening tests, ABX discrimination tests, speaker similarity scoring with Resemblyzer, and acoustic analyses based on mel-cepstral distortion (MCD) and F0 RMSE over 960 audio pairs. Results reveal substantial variation in TTS performance across speech domains, with emotional speech consistently presenting the greatest synthesis challenge (mean MCD 12.03 dB; mean F0 RMSE 889 cents), while conversational speech achieves the highest overall acoustic fidelity. Beyond the empirical findings, this work provides a reproducible evaluation framework, publicly releasing evaluation scripts, result tables, and executable Colab notebooks to support standardized benchmarking and future research on TTS evaluation for low-resource languages.
Problem

Research questions and friction points this paper is trying to address.

text-to-speech evaluation
low-resource languages
domain-specific analysis
speech synthesis benchmarking
acoustic fidelity
Innovation

Methods, ideas, or system contributions that make the work stand out.

domain-specific evaluation
multi-metric benchmarking
low-resource TTS
acoustic fidelity
reproducible framework
🔎 Similar Papers
A
Ali Jafar
FAST School of Computing, National University of Computer and Emerging Sciences (FAST-NUCES), Lahore, Punjab, Pakistan
A
Amal Sarmad
FAST School of Computing, National University of Computer and Emerging Sciences (FAST-NUCES), Lahore, Punjab, Pakistan
S
Shifa Yousaf
FAST School of Computing, National University of Computer and Emerging Sciences (FAST-NUCES), Lahore, Punjab, Pakistan
M
Maryam Bashir
FAST School of Computing, National University of Computer and Emerging Sciences (FAST-NUCES), Lahore, Punjab, Pakistan