Institution profile

Giant Network

Industry researchasia · cn
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

UniVoice: A Unified Model for Speech and Singing Voice Generation

Jun 04, 2026

Traditional approaches struggle to jointly model natural speech and controllable singing within a single framework due to the fundamentally different constraints imposed by prosody and melody. This work proposes UniVoice, a unified generative architecture based on conditional flow matching that decomposes conditioning signals into content, melody, and timbre, and employs a shared DiT backbone for multimodal fusion. A key innovation is the introduction of learnable null melody tokens to replace explicit melody inputs during speech synthesis, thereby preserving explicit melodic control for singing while avoiding redundant constraints on speech; theoretical analysis shows this design approximates marginalization over melody. Trained on 30k hours of speech and 35k hours of singing data, UniVoice achieves a speech PER of 5.26%, comparable to specialized TTS systems, and a singing PER of 16.22%, substantially outperforming the unified baseline Vevo1.5 (24.72%).

0 citationsRead paper

YingVideo-MV: Music-Driven Multi-Stage Video Generation

Dec 02, 2025

Existing music-driven virtual human video generation methods struggle to model camera motion and long-term temporal coherence. To address this, we propose the first music-video generation framework with explicit camera control, featuring a multi-stage cascaded architecture: (1) the MV-Director module enables interpretable shot planning and synchronized audio-motion-camera alignment; (2) a temporally aware diffusion Transformer captures long-range spatiotemporal dependencies; and (3) a latent-space camera adapter combined with audio-embedding-guided dynamic denoising enhances motion naturalness. To support training, we introduce Music-in-the-Wild, a large-scale, diverse dataset of in-the-wild musical performances. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches across multiple benchmarks, enabling synthesis of high-fidelity, highly coherent music performance videos lasting several minutes—complete with realistic, controllable camera motion.

0 citationsRead paper

R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion

Oct 23, 2025

To address the dual challenges of environmental noise interference and inadequate expressiveness modeling in real-world singing voice conversion (SVC), this paper proposes R2-SVC—a robust, expressive SVC framework. First, it constructs a noise-robust training dataset by simulating realistic degradations—including music separation artifacts and random fundamental frequency perturbations—and filtering samples via DNSMOS. Second, it integrates domain-adapted speaker representation learning with a neural source-filter (NSF) architecture to explicitly disentangle harmonic and noise components, thereby enhancing timbral controllability and naturalness. Evaluated on a multi-noise-condition SVC benchmark, R2-SVC achieves state-of-the-art performance, significantly improving conversion quality, inference stability, and cross-condition generalization. Notably, it is the first work to systematically bridge the gap between clean-data training and noisy real-world inference, advancing practical SVC deployment.

0 citationsRead paper

DeepDubber-V1: Towards High Quality and Dialogue, Narration, Monologue Adaptive Movie Dubbing Via Multi-Modal Chain-of-Thoughts Reasoning Guidance

Mar 31, 2025

To address insufficient dialogue/narration/monologue style adaptation and weak modeling of fine-grained speaker attributes (e.g., age, gender) and contextual styles in film dubbing, this paper proposes a vision-driven multimodal large language model framework, introducing for the first time a chain-of-thought (CoT)-guided dubbing style perception and generation paradigm. Our method integrates vision–speech cross-modal CoT reasoning, a conditional large speech synthesis model (variants of VITS/Grad-TTS), and fine-grained speaker representation learning. Key contributions include: (1) releasing CoT-Movie-Dubbing—the first film dubbing dataset with CoT annotations; and (2) achieving breakthroughs in dubbing-type adaptation and joint emotion–timbre modeling. Extensive experiments on V2C, Grid, and CoT-Movie-Dubbing demonstrate state-of-the-art performance: +19.39% SPK-SIM, +12.64% EMO-SIM, −29.49 percentage points WER reduction, and significant improvements in LSE-D and MCD-SL.

0 citationsRead paper

DeepSound-V1: Start to Think Step-by-Step in the Audio Generation from Videos

Mar 28, 2025

Existing video-to-audio generation methods suffer from imprecise visual-audio temporal and semantic alignment, primarily due to the lack of fine-grained alignment annotations in open benchmark datasets. To address this, we propose a novel inference-guided generation framework that requires no manual annotation: (1) we pioneer the integration of chain-of-thought (CoT) reasoning from multimodal large language models (MLLMs) directly into the audio synthesis pipeline, enabling stepwise, interpretable cross-modal alignment; (2) we construct the first video-audio-text multimodal reasoning dataset explicitly designed for inference-guided generation; and (3) we jointly optimize self-supervised temporal reasoning and cross-modal alignment modeling. Experiments demonstrate significant improvements: a 0.89% reduction in voice misalignment rate, and decreases of 10.07% in FDPaSST, 11.62% in FDPA-NNs, and 38.61% in FDVGG—alongside gains of 4.95% in Inception Score (IS) and 6.39% in IB-score—establishing new state-of-the-art performance.

0 citationsRead paper
Recent publications

Latest Papers

UniVoice: A Unified Model for Speech and Singing Voice Generation

Jun 04, 2026

Traditional approaches struggle to jointly model natural speech and controllable singing within a single framework due to the fundamentally different constraints imposed by prosody and melody. This work proposes UniVoice, a unified generative architecture based on conditional flow matching that decomposes conditioning signals into content, melody, and timbre, and employs a shared DiT backbone for multimodal fusion. A key innovation is the introduction of learnable null melody tokens to replace explicit melody inputs during speech synthesis, thereby preserving explicit melodic control for singing while avoiding redundant constraints on speech; theoretical analysis shows this design approximates marginalization over melody. Trained on 30k hours of speech and 35k hours of singing data, UniVoice achieves a speech PER of 5.26%, comparable to specialized TTS systems, and a singing PER of 16.22%, substantially outperforming the unified baseline Vevo1.5 (24.72%).

0 citationsRead paper

YingVideo-MV: Music-Driven Multi-Stage Video Generation

Dec 02, 2025

Existing music-driven virtual human video generation methods struggle to model camera motion and long-term temporal coherence. To address this, we propose the first music-video generation framework with explicit camera control, featuring a multi-stage cascaded architecture: (1) the MV-Director module enables interpretable shot planning and synchronized audio-motion-camera alignment; (2) a temporally aware diffusion Transformer captures long-range spatiotemporal dependencies; and (3) a latent-space camera adapter combined with audio-embedding-guided dynamic denoising enhances motion naturalness. To support training, we introduce Music-in-the-Wild, a large-scale, diverse dataset of in-the-wild musical performances. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches across multiple benchmarks, enabling synthesis of high-fidelity, highly coherent music performance videos lasting several minutes—complete with realistic, controllable camera motion.

0 citationsRead paper

R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion

Oct 23, 2025

To address the dual challenges of environmental noise interference and inadequate expressiveness modeling in real-world singing voice conversion (SVC), this paper proposes R2-SVC—a robust, expressive SVC framework. First, it constructs a noise-robust training dataset by simulating realistic degradations—including music separation artifacts and random fundamental frequency perturbations—and filtering samples via DNSMOS. Second, it integrates domain-adapted speaker representation learning with a neural source-filter (NSF) architecture to explicitly disentangle harmonic and noise components, thereby enhancing timbral controllability and naturalness. Evaluated on a multi-noise-condition SVC benchmark, R2-SVC achieves state-of-the-art performance, significantly improving conversion quality, inference stability, and cross-condition generalization. Notably, it is the first work to systematically bridge the gap between clean-data training and noisy real-world inference, advancing practical SVC deployment.

0 citationsRead paper

DeepDubber-V1: Towards High Quality and Dialogue, Narration, Monologue Adaptive Movie Dubbing Via Multi-Modal Chain-of-Thoughts Reasoning Guidance

Mar 31, 2025

To address insufficient dialogue/narration/monologue style adaptation and weak modeling of fine-grained speaker attributes (e.g., age, gender) and contextual styles in film dubbing, this paper proposes a vision-driven multimodal large language model framework, introducing for the first time a chain-of-thought (CoT)-guided dubbing style perception and generation paradigm. Our method integrates vision–speech cross-modal CoT reasoning, a conditional large speech synthesis model (variants of VITS/Grad-TTS), and fine-grained speaker representation learning. Key contributions include: (1) releasing CoT-Movie-Dubbing—the first film dubbing dataset with CoT annotations; and (2) achieving breakthroughs in dubbing-type adaptation and joint emotion–timbre modeling. Extensive experiments on V2C, Grid, and CoT-Movie-Dubbing demonstrate state-of-the-art performance: +19.39% SPK-SIM, +12.64% EMO-SIM, −29.49 percentage points WER reduction, and significant improvements in LSE-D and MCD-SL.

0 citationsRead paper

DeepSound-V1: Start to Think Step-by-Step in the Audio Generation from Videos

Mar 28, 2025

Existing video-to-audio generation methods suffer from imprecise visual-audio temporal and semantic alignment, primarily due to the lack of fine-grained alignment annotations in open benchmark datasets. To address this, we propose a novel inference-guided generation framework that requires no manual annotation: (1) we pioneer the integration of chain-of-thought (CoT) reasoning from multimodal large language models (MLLMs) directly into the audio synthesis pipeline, enabling stepwise, interpretable cross-modal alignment; (2) we construct the first video-audio-text multimodal reasoning dataset explicitly designed for inference-guided generation; and (3) we jointly optimize self-supervised temporal reasoning and cross-modal alignment modeling. Experiments demonstrate significant improvements: a 0.89% reduction in voice misalignment rate, and decreases of 10.07% in FDPaSST, 11.62% in FDPA-NNs, and 38.61% in FDVGG—alongside gains of 4.95% in Inception Score (IS) and 6.39% in IB-score—establishing new state-of-the-art performance.

0 citationsRead paper