Institution profile

Migu

Industry researchasia · cn
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment

Oct 10, 2025

Data scarcity and limited model scalability hinder the development of diffusion-based singing synthesis. To address these challenges, this work proposes a two-stage solution: First, a high-quality Chinese singing dataset exceeding 500 hours is constructed by leveraging large language models (LLMs) to generate diverse lyrics aligned with fixed melodies. Second, DiTSinger—a novel diffusion singing synthesizer—is introduced, integrating a Diffusion Transformer architecture with an implicit alignment mechanism that eliminates reliance on phoneme-level duration annotations; instead, character-level speech attention constraints enhance alignment robustness. Furthermore, the model’s depth, width, and feature resolution are systematically scaled to improve representational capacity. Experiments demonstrate stable training and high-fidelity synthesis even without precise alignment labels, achieving significant improvements over existing diffusion-based methods in scalability, robustness, and audio quality.

0 citationsRead paper
Recent publications

Latest Papers

DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment

Oct 10, 2025

Data scarcity and limited model scalability hinder the development of diffusion-based singing synthesis. To address these challenges, this work proposes a two-stage solution: First, a high-quality Chinese singing dataset exceeding 500 hours is constructed by leveraging large language models (LLMs) to generate diverse lyrics aligned with fixed melodies. Second, DiTSinger—a novel diffusion singing synthesizer—is introduced, integrating a Diffusion Transformer architecture with an implicit alignment mechanism that eliminates reliance on phoneme-level duration annotations; instead, character-level speech attention constraints enhance alignment robustness. Furthermore, the model’s depth, width, and feature resolution are systematically scaled to improve representational capacity. Experiments demonstrate stable training and high-fidelity synthesis even without precise alignment labels, achieving significant improvements over existing diffusion-based methods in scalability, robustness, and audio quality.

0 citationsRead paper