The VoiceMOS Challenge 2026: Evaluating Speech Enhancement, Emotional TTS and Accented TTS Systems

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过组织三个不同赛道,评估了语音增强、情感TTS及口音TTS系统的主观语音质量自动预测方法,吸引了18个团队参与并超越基线。
📝 Abstract
We present the results of the VoiceMOS Challenge 2026, the fifth edition of a scientific challenge on automatic prediction of subjective speech assessments. After expanding the scope to music and general audio in 2025, we refocused the evaluation target on speech and organized three tracks: prediction of absolute and comparative category ratings for enhanced speech, prediction of naturalness and emotion-related tasks for emotional text-to-speech systems, and prediction of speaker and accent similarity for codec-based speech synthesis systems. The challenge attracted a total of 18 teams worldwide, with most teams successfully surpassing the provided baselines. We summarize the challenge results, representative top-performing systems, participant feedback, and directions for future editions.
Problem

Research questions and friction points this paper is trying to address.

Speech Enhancement
Emotional TTS
Accented TTS
Innovation

Methods, ideas, or system contributions that make the work stand out.

speech enhancement
emotional TTS
accented TTS