Is Semantics Enough for Speech Mean Opinion Score Prediction?

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了语义与声学细节对语音质量评估的影响,通过对比自监督学习模型、纯声学神经音频编码器及结合两者的模型,证明了结合语义理解和精细声学建模能更准确预测语音平均意见分。
📝 Abstract
Mean Opinion Score (MOS) is the gold standard for evaluating synthesized speech naturalness. However, current automatic MOS predictors are dominated by self-supervised learning (SSL) models that prioritize high-level semantics, potentially compromising their ability to capture critical acoustic details. In this paper, we systematically investigate representations from three paradigms: SSLs, acoustic-only neural audio codecs (NACs), and unified NACs that integrate semantics into reconstruction-based architectures. Extensive benchmarking on the standard BVCC and multiple out-of-domain (OOD) datasets demonstrates that features synergizing semantic understanding with fine-grained acoustic modeling achieve a higher performance upper bound in speech quality assessment. Ultimately, our findings highlight that semantics alone are not enough; a dual focus on semantic content and acoustic fidelity is essential for robust MOS prediction.
Problem

Research questions and friction points this paper is trying to address.

Mean Opinion Score
self-supervised learning
acoustic details
speech quality assessment
semantic understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-supervised learning
neural audio codecs
semantic and acoustic integration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tianyu Lan
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China
Yufei Shi
Yufei Shi
National University of Singapore
Vision computing
Yang Ai
Yang Ai
Associate Researcher, University of Science and Technology of China
Speech SynthesisSpeech EnhancementSpeech CodingDeep Learning
H
Honghao Sun
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China
H
Huipeng Du
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China
Z
Zhenhua Ling
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China