StepAudio 3 Gen Technical Report

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出StepAudio 3 Gen,一种基于残差向量量化(RVQ)标记的离散自回归生成模型,用于解决多种音频类型生成问题,包括文本转语音、声音设计等。
📝 Abstract
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization (RVQ) tokens, departing from the diffusion Transformer-based continuous generation paradigm prevalent in recent general audio models. Its StepAudio Tokenizer represents general audio at 12.5 Hz in a shared $16 \times 2048$ residual code space, jointly quantizing semantic and waveform-level acoustic features so that each code layer preserves both types of information. For generation, the backbone predicts the first codebook along the time axis using autoregressive modeling, while a lightweight causal Transformer completes the remaining fifteen codebooks along the codebook axis. Our study further identifies three key design principles: (1) interference-aware progressive pretraining for acquiring audio capabilities while preserving the textual abilities of the large language model, (2) RVQ Adaptor for effectively incorporating multi-codebook acoustic representations, and (3) discrete autoregressive modeling over a shared representation across general audio domains. With progressive pretraining, multi-task instruction training, and supervised fine-tuning, StepAudio 3 Gen achieves state-of-the-art performance on both TTS and voice design, while retaining strong generation capabilities across speech, vocals, sound effects, and music. Audio samples are available at https://stepaudiollm.github.io/step-audio-3-gen/.
Problem

Research questions and friction points this paper is trying to address.

audio generation
text-to-speech
voice design
vocal generation
sound effects
Innovation

Methods, ideas, or system contributions that make the work stand out.

discrete autoregressive generator
residual vector quantization (RVQ)
shared residual code space
interference-aware progressive pretraining
RVQ Adaptor
🔎 Similar Papers
2024-04-22arXiv.orgCitations: 25
B
Bin Lin
B
Bo Zhao
B
Boyang Wang
B
Boyang Zhang
B
Boyong Wu
Chao Yan
Chao Yan
Instructor at DBMI, VUMC; CS PhD from Vanderbilt U
AI for medicineSynthetic health dataPrivacyFairness
Chen Geng
Chen Geng
Stanford University
4D VisionComputer GraphicsInverse Graphics
C
Chen Wu
C
Cheng Yi
C
Chengli Feng
C
Chenglin Zhu
D
DanNi Wan
Daxin Jiang
Daxin Jiang
Co-Founder & CEO, StepFun Corporation
Deep LearningFoundation Models
D
Dongqing Pang
F
Fei Tian
F
Feng Tian
F
Future Li
G
Gang Yu
G
Guanglong Yang
J
Jia Peng
J
Jiahao Song
J
Jiamin Fan
J
Jiangjie Zhen
J
Jianzheng Gao
J
Jun Chen