🤖 AI Summary
This study addresses the limitations of existing automatic TTS evaluation methods, which overly rely on coarse-grained “naturalness” scores and fail to capture the multidimensional nature of human speech quality perception. The authors decompose “naturalness” into ten linguistically grounded, fine-grained dimensions and introduce the first dimension-level TTS meta-evaluation benchmark, comprising 860 expert-annotated samples. They systematically evaluate both MOS predictors and Audio-LLMs against this benchmark, revealing that MOS predictors primarily reflect acoustic quality, while Audio-LLMs exhibit prompt-dependent performance and limited generalization. Critically, neither approach reliably detects structured linguistic errors in synthetic speech. The work releases its dataset and code to advance interpretable, multidimensional TTS evaluation research.
📝 Abstract
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.