DualSpecSE: A Dual-Path Speech Enhancement Network Integrating Mel and Complex Spectrograms

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出DualSpecSE,通过结合Mel谱图和复数谱图的双路径架构来提升语音增强效果,提高自动语音识别性能和语音重建质量。
📝 Abstract
In this paper, we propose DualSpecSE, a speech enhancement framework that jointly models Mel-spectrogram and complex spectrogram in a dual-path architecture for improved ASR performance and higher-quality speech reconstruction. The Mel branch learns coarse-grained acoustic representations and produces enhanced Mel-spectrograms for direct ASR usage, while the complex branch refines fine-grained spectral details for high-fidelity waveform reconstruction. Built upon the cross-band and narrow-band blocks from CleanMel, DualSpecSE introduces an interaction module and a fusion module to enable effective information exchange between the two branches. The model simultaneously outputs enhanced Mel and complex spectrogram without requiring a pretrained vocoder. Experimental results demonstrate consistent improvements in speech fidelity, perceptual quality, and ASR performance. Codes and audio samples are available.
Problem

Research questions and friction points this paper is trying to address.

speech enhancement
ASR performance
speech reconstruction
Mel-spectrogram
complex spectrogram
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-Path Architecture
Interaction Module
Fusion Module
Mel-spectrogram and Complex Spectrogram
💼 Related Jobs
No related jobs found.
X
Xingchen Li
Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, China
Ziqian Wang
Ziqian Wang
Northwestern Polytechnical University
speech processingdeep learning
Z
Zikai Liu
Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, China
Y
Yike Zhu
Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, China
Z
Zihan Zhang
Huawei Technologies Co., Ltd., China
L
Longshuai Xiao
Huawei Technologies Co., Ltd., China
Lei Xie
Lei Xie
Northwestern Polytechnical University
speech processingspeech recognitionspeech synthesismultimediaartificial intelligence