DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment
Data scarcity and limited model scalability hinder the development of diffusion-based singing synthesis. To address these challenges, this work proposes a two-stage solution: First, a high-quality Chinese singing dataset exceeding 500 hours is constructed by leveraging large language models (LLMs) to generate diverse lyrics aligned with fixed melodies. Second, DiTSinger—a novel diffusion singing synthesizer—is introduced, integrating a Diffusion Transformer architecture with an implicit alignment mechanism that eliminates reliance on phoneme-level duration annotations; instead, character-level speech attention constraints enhance alignment robustness. Furthermore, the model’s depth, width, and feature resolution are systematically scaled to improve representational capacity. Experiments demonstrate stable training and high-fidelity synthesis even without precise alignment labels, achieving significant improvements over existing diffusion-based methods in scalability, robustness, and audio quality.