Is Prosody Lost in Translation? Fine-Grained Cross-Lingual Prosody Similarity Across Languages

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过分析英语与德语、西班牙语、法语之间的音高、能量和时序特征相似性,探讨了跨语言韵律模式的异同,为表达性语音翻译系统提供了指导。
📝 Abstract
Prosody plays an important role in speech translation, conveying information such as emphasis, emotion, and intent beyond lexical content. However, despite recent progress in expressive speech-to-speech translation (S2ST), little is known about how prosodic patterns are similar/different across languages. Understanding these cross-lingual similarities and differences is crucial for effectively incorporating prosody into expressive S2ST systems. In this work, we present the first fine-grained cross-lingual analysis of prosody using multilingual dubbing data across English-German, English-Spanish, and English-French language pairs. We analyze the similarity of pitch, energy, and temporal feature patterns between source and target speech and investigate the linguistic and alignment-related factors affecting this similarity. Our analysis reveals inherent cross-lingual correlations in prosodic structure between certain languages. The findings provide important insights into the transferability of prosody across languages and offer empirical guidance for future expressive speech-to-speech translation systems.
Problem

Research questions and friction points this paper is trying to address.

Prosody
Cross-lingual
Speech-to-speech Translation
Expressive S2ST
Multilingual Dubbing
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-lingual prosody
fine-grained analysis
expressive S2ST
multilingual dubbing data
prosodic structure
🔎 Similar Papers
2023-10-10arXiv.orgCitations: 0
💼 Related Jobs
No related jobs found.
H
Haopeng Xie
Center for Language and Speech Processing (CLSP), Johns Hopkins University
I
Ismail Rasim Ulgen
Center for Language and Speech Processing (CLSP), Johns Hopkins University
S
Sofia Son
Center for Language and Speech Processing (CLSP), Johns Hopkins University
Berrak Sisman
Berrak Sisman
Assistant Professor (ECE & DSAI), Johns Hopkins University
Machine LearningAffective ComputingSpeech SynthesisVoice ConversionAnti-spoofing
Philipp Koehn
Philipp Koehn
Professor, Johns Hopkins University
Machine TranslationNatural Language Processing