Preference Optimization for Non-Verbal Vocalization Synthesis

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过优化偏好信号、构建偏好对及使用DPO目标,解决非言语发声合成的有效性问题,提出NV-CER评估方法以改进TTS中的非言语发声实现。
📝 Abstract
Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive TTS, but the effectiveness of preference optimization for NV generation remains poorly understood. We systematically study preference optimization for NV-capable TTS, focusing on preference signals, preference-pair construction, and DPO-based optimization objectives. We formulate an NV-aware character error rate (NV-CER) by treating NV tags as distinct output symbols and computing a weighted pinyin-based CER over both verbal and non-verbal content, enabling controllable optimization of NV realization without modifying the underlying optimization algorithm. Experiments on Emilia-NV and the augmented NV-Bench covering 18 NV types reveal how different design choices affect NV realization and lexical fidelity, and establish an effective setup using standard DPO. Objective, LLM-based, and human evaluations provide converging evidence for our findings, offering practical insights into NV-aware post-training for expressive TTS.
Problem

Research questions and friction points this paper is trying to address.

Non-verbal vocalizations
Preference optimization
TTS
NV-aware
DPO
Innovation

Methods, ideas, or system contributions that make the work stand out.

Preference Optimization
Non-Verbal Vocalizations
NV-CER
DPO-based Optimization
🔎 Similar Papers