Institution profile

Ximalaya Inc.

Industry researchasia · cn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

S2ST-Omni: An Efficient and Scalable Multilingual Speech-to-Speech Translation Framework via Seamlessly Speech-Text Alignment and Streaming Speech Decoder

Jun 11, 2025

Multilingual speech-to-speech translation (S2ST) faces two major challenges: the trade-off between high translation quality and low end-to-end latency, and heavy reliance on scarce parallel speech data. To address these, we propose an end-to-end decoupled framework that jointly models speech-to-text (S2TT) and text-to-speech (TTS) components. It employs a lightweight speech adapter to align cross-modal representations, integrates Whisper’s audio encoder with Qwen-3.0’s strong textual understanding, and introduces a streaming autoregressive TTS decoder to ensure real-time inference. Our approach drastically reduces dependence on parallel speech corpora while achieving state-of-the-art BLEU and COMET scores on the CVSS benchmark—outperforming existing S2ST systems. Crucially, its end-to-end latency matches that of the best-performing baselines, demonstrating strong practical deployability for real-world multilingual translation applications.

0 citationsRead paper

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech

May 20, 2025

This paper addresses high-fidelity, fine-grained controllable emotional voice conversion. We propose a novel method enabling dual-path control—via natural language prompts or reference speech—and continuous adjustment of emotional intensity. Our key contributions are: (1) the first emotion-aware contrastive language–audio pretraining model, EVC-CLAP, which enhances cross-modal emotional alignment; (2) FuEncoder, a speaker- and emotion-encoding module with adaptive intensity gating, enabling disentangled and controllable emotion intensity representation; and (3) an end-to-end flow-matching-based reconstruction framework integrating Phonetic PosteriorGrams and ASR-derived auxiliary representations. Comprehensive objective and subjective evaluations demonstrate state-of-the-art performance: MOS of 4.12, with significant improvements in emotion accuracy, naturalness, and controllability over prior methods.

0 citationsRead paper
Recent publications

Latest Papers

S2ST-Omni: An Efficient and Scalable Multilingual Speech-to-Speech Translation Framework via Seamlessly Speech-Text Alignment and Streaming Speech Decoder

Jun 11, 2025

Multilingual speech-to-speech translation (S2ST) faces two major challenges: the trade-off between high translation quality and low end-to-end latency, and heavy reliance on scarce parallel speech data. To address these, we propose an end-to-end decoupled framework that jointly models speech-to-text (S2TT) and text-to-speech (TTS) components. It employs a lightweight speech adapter to align cross-modal representations, integrates Whisper’s audio encoder with Qwen-3.0’s strong textual understanding, and introduces a streaming autoregressive TTS decoder to ensure real-time inference. Our approach drastically reduces dependence on parallel speech corpora while achieving state-of-the-art BLEU and COMET scores on the CVSS benchmark—outperforming existing S2ST systems. Crucially, its end-to-end latency matches that of the best-performing baselines, demonstrating strong practical deployability for real-world multilingual translation applications.

0 citationsRead paper

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech

May 20, 2025

This paper addresses high-fidelity, fine-grained controllable emotional voice conversion. We propose a novel method enabling dual-path control—via natural language prompts or reference speech—and continuous adjustment of emotional intensity. Our key contributions are: (1) the first emotion-aware contrastive language–audio pretraining model, EVC-CLAP, which enhances cross-modal emotional alignment; (2) FuEncoder, a speaker- and emotion-encoding module with adaptive intensity gating, enabling disentangled and controllable emotion intensity representation; and (3) an end-to-end flow-matching-based reconstruction framework integrating Phonetic PosteriorGrams and ASR-derived auxiliary representations. Comprehensive objective and subjective evaluations demonstrate state-of-the-art performance: MOS of 4.12, with significant improvements in emotion accuracy, naturalness, and controllability over prior methods.

0 citationsRead paper