PACodec: A Low-bitrate Neural Speech Codec with Parallel Additive Vector Quantization

๐Ÿ“… 2026-09-03
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บPACodec๏ผŒไธ€็งๅŸบไบŽๅนถ่กŒๅŠ ๆ€ง็Ÿข้‡้‡ๅŒ–็š„ๆ–ฐไฝŽๆฏ”็‰น็އ็ฅž็ป่ฏญ้Ÿณ็ผ–็ ๅ™จ๏ผŒ้€š่ฟ‡ไผ˜ๅŒ–็ ็އไฝฟ็”จๅ‡ๅฐ‘30%ๆฏ”็‰น็އ๏ผŒๅŒๆ—ถไฟๆŒ่งฃ็ ่ดจ้‡ใ€‚
๐Ÿ“ Abstract
This paper proposes PACodec, a novel low-bitrate neural speech codec based on parallel additive vector quantization (PAVQ). Unlike the mainstream residual vector quantization (RVQ) used in most neural speech codecs, where vector quantizers (VQs) are sequentially dependent, the PAVQ strategy adopted in PACodec aggregates parallel quantization results to optimize bitrate usage. Specifically, the PAVQ adopts a "global-local-global" (GLG) design: the global encoded features are quantized in parallel by multiple independent VQs, each attending to a local component of the representation, and their outputs are aggregated through addition to yield the final global quantization result for decoding. Experimental results show that PACodec, as each VQ focuses only on local information, supports smaller codebooks and reduces bitrate by 30% compared with baselines at the same decoding quality, with only minor model complexity. Further analysis shows that, owing to the GLG framework of PAVQ, the proposed PACodec is disentanglement-friendly, and each independent VQ captures different aspects of speech, e.g., content, timbre, and acoustic details, suggesting potential for application to downstream tasks such as voice conversion.
Problem

Research questions and friction points this paper is trying to address.

low-bitrate
neural speech codec
vector quantization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parallel Additive Vector Quantization
Low-bitrate Neural Speech Codec
Global-Local-Global Design
Bitrate Reduction
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
F
Fei Liu
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China
Yang Ai
Yang Ai
Associate Researcher, University of Science and Technology of China
Speech SynthesisSpeech EnhancementSpeech CodingDeep Learning
X
Xiao-Hang Jiang
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China
Z
Zhen-Hua Ling
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China