Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对低资源希腊语TTS质量问题,通过数据处理和模型微调方法,利用确定性提示解决了说话人一致性问题。
📝 Abstract
Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, lacking the curated corpora behind state-of-the-art synthesis. We propose a data curation recipe that transforms audiobook recordings into TTS-ready data via WhisperX alignment and filtering. Then we fine-tune Parler-TTS (880M), a prompt-based multilingual model whose pre-training encodes phonetic priors transferable to Greek. During development, we find that LLM-generated style prompts introduce speaker drift at inference. Replacing them with deterministic prompts resolves this, and a speaker-specific LoRA stage trained on 3.5 h of single-speaker data anchors identity while updating ~5% of parameters. Our system achieves WER 10.7% (2.9 above the ASR floor), MOS-I 4.00 (vs. 4.36 human speech), and near-human speaker consistency (MOS-C 4.24 vs. 4.30), showing that robust single-speaker Greek TTS is achievable with limited curated data.
Problem

Research questions and friction points this paper is trying to address.

Low-Resource TTS
Greek
Data Scarcity
Speaker Stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

deterministic prompting
speaker drift
LoRA
low-resource TTS
WhisperX
🔎 Similar Papers
No similar papers found.