ProLombard: Structured Multi-Scale Modeling for Normal-to-Lombard Speech Conversion

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出ProLombard框架,通过多尺度建模和对齐说话人编码器解决正常语音到朗巴德语音转换中语音清晰度、说话人身份保持及内容分离问题。
📝 Abstract
Normal-to-Lombard (N2L) speech conversion aims to improve speech intelligibility in noisy environments by transforming normal speech into Lombard-style speech while preserving linguistic content, speaker identity, and speech quality. Despite recent progress, existing methods typically model the Lombard effect at the utterance level or the frame level, overlooking its hierarchical nature and its entanglement with both speaker identity and phoneme-level content. This limitation leads to Lombard leakage in speaker representations and incomplete separation between Lombard characteristics and linguistic content. In this work, we propose ProLombard, a structured multi-scale N2L framework that explicitly models the Lombard effect across utterance-, phoneme-, and frame-level representations. To address Lombard-speaker entanglement, we introduce an aligned speaker encoder (ASE) that suppresses Lombard leakage by aligning Lombard-speech speaker embeddings with their normal-speech counterparts. To achieve more complete Lombard-content disentanglement, we develop a phoneme-aware disentanglement and injection mechanism that extends conventional frame-level modeling to the phoneme level. Furthermore, we design a vector quantization (VQ)-median module that provides robust phoneme-level representations through VQ-based segmentation and median-frame-based aggregation. Extensive experiments on Mandarin and English Lombard datasets demonstrate that the proposed approach consistently improves speech intelligibility, Lombard similarity, and perceptual quality over baselines while maintaining speaker identity. These results highlight the importance of structured multi-scale modeling for effective N2L speech conversion.
Problem

Research questions and friction points this paper is trying to address.

Normal-to-Lombard speech conversion
Lombard effect
speaker identity
phoneme-level content
Lombard leakage
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured Multi-Scale Modeling
Aligned Speaker Encoder (ASE)
Phoneme-Aware Disentanglement and Injection
Vector Quantization (VQ)-median Module
Hongyang Chen
Hongyang Chen
SUN YAT-SEN UNIVERSITY
SDNCloud ComputingMicroserviceAIOps
X
Xinmeng Xu
National Engineering Research Center for Multimedia Software (NERCMS), School of Computer Science, Wuhan University, Wuhan 430072, China; and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University, Wuhan 430072, China
Y
Youqiang Zheng
National Engineering Research Center for Multimedia Software (NERCMS), School of Computer Science, Wuhan University, Wuhan 430072, China; and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University, Wuhan 430072, China
X
Xingyu Liu
National Engineering Research Center for Multimedia Software (NERCMS), School of Computer Science, Wuhan University, Wuhan 430072, China; and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University, Wuhan 430072, China
Y
Yuhong Yang
National Engineering Research Center for Multimedia Software (NERCMS), School of Computer Science, Wuhan University, Wuhan 430072, China; and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University, Wuhan 430072, China
Zhongyuan Wang
Zhongyuan Wang
Wuhan University
Weiping Tu
Weiping Tu
Wuhan University, Wuhan City, Hubei Prov., China
audio signal processingartificial intelligence
S
Song Lin
Guangdong OPPO Mobile Telecommunications Corp., China