PhaseGAN: High-Fidelity Vocoder via Decoupled Amplitude and GAN-Driven Phase Reconstruction

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
PhaseGAN通过解耦幅度和基于GAN的相位重建方法,解决了神经网络声码器中准确相位重建的问题,提高了音频质量和建模效率。
📝 Abstract
A vocoder is a pivotal component of modern text-to-speech (TTS) systems. Despite the significant progress of neural network-based vocoders, accurate phase reconstruction remains the main challenge limiting both audio quality and modeling efficiency. We introduce PhaseGAN, a lightweight vocoder that addresses this limitation through a "mel $\rightarrow$ Amplitude $\rightarrow$ Phase" reconstruction pipeline. By reconstructing amplitude and phase spectra via distinct methodologies, the proposed PhaseGAN outperforms state-of-the-art baselines while utilizing fewer model parameters and reduced computational requirements. The compact version generates high-fidelity audio with approximately 500K parameters and 1 GMAC computational load, making it highly suitable for real-time applications on edge devices. In addition, our approach exhibits exceptional musical audio synthesis capabilities despite no training on musical data, illustrating unprecedented cross-domain generalization. See https://github.com/phasegan/phasegan-audio-demo for demos of our work.
Problem

Research questions and friction points this paper is trying to address.

vocoder
phase reconstruction
audio quality
modeling efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

PhaseGAN
decoupled amplitude and phase reconstruction
lightweight vocoder
high-fidelity audio
cross-domain generalization
🔎 Similar Papers
No similar papers found.
Wenzheng Zhang
Wenzheng Zhang
Rutgers University
Natural Language ProcessingDeep Learning
Xueliang Zhang
Xueliang Zhang
Inner Mongolia University
Speech enhancementSpeech separationComputational Auditory Scene Analysis
S
Shulin He
College of Computer Science, Inner Mongolia University, China
Fei Zhao
Fei Zhao
School of Computer Science, Inner Mongolia University
Acoustic Echo CancellationActive Noise Control
X
Xin Liu
College of Computer Science, Inner Mongolia University, China
P
Pengjie Shen
College of Computer Science, Inner Mongolia University, China
Z
Zhenlong Guo
College of Computer Science, Inner Mongolia University, China
Z
Zixuan Xue
College of Computer Science, Inner Mongolia University, China
H
Hongtao Bao
College of Computer Science, Inner Mongolia University, China
Zixuan Li
Zixuan Li
Assistant Professor at ICT, UCAS
Knowledge GraphLarge Language Model