Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing automatic TTS evaluation methods, which overly rely on coarse-grained “naturalness” scores and fail to capture the multidimensional nature of human speech quality perception. The authors decompose “naturalness” into ten linguistically grounded, fine-grained dimensions and introduce the first dimension-level TTS meta-evaluation benchmark, comprising 860 expert-annotated samples. They systematically evaluate both MOS predictors and Audio-LLMs against this benchmark, revealing that MOS predictors primarily reflect acoustic quality, while Audio-LLMs exhibit prompt-dependent performance and limited generalization. Critically, neither approach reliably detects structured linguistic errors in synthetic speech. The work releases its dataset and code to advance interpretable, multidimensional TTS evaluation research.
📝 Abstract
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.
Problem

Research questions and friction points this paper is trying to address.

Text-to-Speech evaluation
naturalness
perceptual dimensions
linguistically grounded
automatic speech assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

linguistically grounded evaluation
TTS naturalness deconstruction
dimension-level benchmark
Audio-LLM judges
MOS predictors
🔎 Similar Papers
2023-10-10arXiv.orgCitations: 0
💼 Related Jobs
No related jobs found.
Oluwanifemi Bamgbose
Oluwanifemi Bamgbose
University of Waterloo
S
Simon Rosen
ServiceNow
J
Jash Shah
ServiceNow
L
Lindsay Devon Brin
ServiceNow
H
Hoang H Nguyen
ServiceNow
A
Anke Koelzer
ServiceNow
R
Rachel Hansen
ServiceNow
T
Tara Bogavelli
ServiceNow
F
Fanny Riols
ServiceNow