Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于连续神经音频编码表示的掩码自回归语音增强方法(MARSE),通过迭代解码掩码干净语音帧来提高语音质量和可懂度。
📝 Abstract
Most previous work on speech enhancement (SE) based on masked generative modeling relied on discrete token representations of audio signals, obtained using neural audio codecs (NACs). However, a recent study has shown that continuous latent representations of NACs can be advantageous for SE in terms of speech quality and intelligibility. In this work, we propose masked autoregressive SE (MARSE), a method for SE based on iterative decoding of masked clean speech frames using continuous NAC representations of speech. In particular, we investigate a set of different decoding policies, ceteris paribus, that is, using the same DNN (a Conformer model), the same NAC (the DAC codec) and the same training setup. The results show that MARSE enables a flexible trade-off between SE performance and computational cost. Audio examples and code are available online.
Problem

Research questions and friction points this paper is trying to address.

speech enhancement
continuous NAC representations
masked autoregressive SE
computational cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

masked autoregressive speech enhancement
continuous neural audio codec representations
iterative decoding
🔎 Similar Papers
No similar papers found.