SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation
This work addresses the limited content fidelity in existing continuous latent autoregressive speech generation methods, which lack explicit linguistic structural supervision. To remedy this, we propose SemBridge, a novel framework that introduces discrete semantic tokens—used only during training—to directly supervise the states of the autoregressive language model. A semantic-aligned acoustic variational autoencoder (VAE) is employed to construct a structured continuous target space. Notably, the inference pipeline remains fully continuous without any architectural modifications. The proposed approach substantially reduces word and character error rates while preserving high speaker similarity and perceptual quality. Furthermore, it enables zero-shot text-to-speech synthesis and score-conditioned singing voice generation.