Neural Music Enhancement with Dual Time-Frequency Spectral Representations for Prediction and Discrimination

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于双时间-频率谱表示的音乐增强模型DSME,使用STFT生成干净音频并用CQT进行判别,以解决非专业录音中的噪音和混响问题。
📝 Abstract
Non-professional music recordings shared online often suffer from background noise and reverberation, degrading perceived quality and limiting reuse. This paper proposes DSME, a music enhancement model based on dual time-frequency spectral representations. Within a generative adversarial framework, DSME uses short-time Fourier transform (STFT) spectra for generation and constant-Q transform (CQT) spectra for discrimination. Leveraging STFT's fixed window, invertibility, and predictability, the generator estimates clean amplitude-phase spectra from degraded inputs and reconstructs waveforms via inverse STFT. Exploiting CQT's log-frequency, variable-window structure aligned with musical octaves, we design an octave-segmented CQT discriminator. We also introduce a chroma-spectrum loss to emphasize pitch and harmonic consistency. Experiments show DSME outperforms baselines in objective and subjective tests, validating the effectiveness of the dual-spectrum approach.
Problem

Research questions and friction points this paper is trying to address.

background noise
reverberation
music enhancement
non-professional recordings
quality degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual Time-Frequency Spectral Representations
Generative Adversarial Framework
Chroma-Spectrum Loss
🔎 Similar Papers
2024-09-26IEEE International Conference on Acoustics, Speech, and Signal ProcessingCitations: 3
F
Fei Liu
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China
Yang Ai
Yang Ai
Associate Researcher, University of Science and Technology of China
Speech SynthesisSpeech EnhancementSpeech CodingDeep Learning
Z
Zhen-Hua Ling
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China