Institution profile

Huya Inc

Industry researchasia · cn
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Divergence-Augmented Policy Optimization

Jan 25, 2025Neural Information Processing Systems

To address the instability and premature convergence caused by reusing offline data in deep reinforcement learning, this paper proposes a Bregman divergence constraint mechanism grounded in state distribution. Differing from conventional approaches that define Bregman divergence over action probability spaces, our method is the first to formulate it over the space of state distributions induced by policies, thereby establishing a divergence-augmented policy optimization framework. By explicitly constraining the magnitude of policy updates’ impact on the induced state distribution, the approach ensures both safety and efficacy in offline data reuse. Evaluated on the Atari benchmark under data-scarce settings, our method significantly improves training stability and convergence speed, while achieving superior sample efficiency and policy robustness compared to mainstream algorithms including PPO and SAC. These results empirically validate the effectiveness and practicality of regularization at the state-distribution level.

13 citations1 influentialRead paper

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS

Jul 07, 2026

Current large language model (LLM)-driven text-to-speech (TTS) systems struggle to enable explicit, fine-grained control over word-level acoustic attributes—such as duration, energy, pitch, and intonation—limiting their applicability in domains like audiobook narration and video dubbing. To address this, this work proposes WordVoice, a novel framework that introduces a boundary-token mechanism to facilitate explicit “acoustic planning” within LLMs. Leveraging the newly curated WordVoice-5A dataset comprising 4.7 thousand hours of bilingual speech, and integrating a fine-grained acoustic modulation module, WordVoice achieves, for the first time in LLM-based TTS, disentangled control over multiple word-level acoustic dimensions. The approach supports both adaptive prosody planning and manual intervention, significantly enhancing the precision of independent control across five acoustic attributes while preserving zero-shot synthesis stability and improving overall speech naturalness.

0 citationsRead paper

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

Jun 26, 2026

This work addresses the limited emotional expressiveness of large language model (LLM)-driven text-to-speech (TTS) systems during supervised fine-tuning, as well as the information conflict and scale mismatch inherent in existing preference optimization approaches. To overcome these challenges, the authors propose the HPRO framework, which introduces a novel differentiable HD-Emo encoder-decoder to disentangle speech into content and style-preference tokens. HPRO further incorporates a hierarchical progressive reward mechanism that aligns optimization objectives across multiple granularities—frame, word, and sentence levels—effectively bridging the gap between frame-level generation and sentence-level rewards. Experimental results demonstrate that HPRO significantly enhances the emotional expressiveness of synthesized speech while preserving linguistic clarity and semantic integrity.

0 citationsRead paper
Recent publications

Latest Papers

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS

Jul 07, 2026

Current large language model (LLM)-driven text-to-speech (TTS) systems struggle to enable explicit, fine-grained control over word-level acoustic attributes—such as duration, energy, pitch, and intonation—limiting their applicability in domains like audiobook narration and video dubbing. To address this, this work proposes WordVoice, a novel framework that introduces a boundary-token mechanism to facilitate explicit “acoustic planning” within LLMs. Leveraging the newly curated WordVoice-5A dataset comprising 4.7 thousand hours of bilingual speech, and integrating a fine-grained acoustic modulation module, WordVoice achieves, for the first time in LLM-based TTS, disentangled control over multiple word-level acoustic dimensions. The approach supports both adaptive prosody planning and manual intervention, significantly enhancing the precision of independent control across five acoustic attributes while preserving zero-shot synthesis stability and improving overall speech naturalness.

0 citationsRead paper

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

Jun 26, 2026

This work addresses the limited emotional expressiveness of large language model (LLM)-driven text-to-speech (TTS) systems during supervised fine-tuning, as well as the information conflict and scale mismatch inherent in existing preference optimization approaches. To overcome these challenges, the authors propose the HPRO framework, which introduces a novel differentiable HD-Emo encoder-decoder to disentangle speech into content and style-preference tokens. HPRO further incorporates a hierarchical progressive reward mechanism that aligns optimization objectives across multiple granularities—frame, word, and sentence levels—effectively bridging the gap between frame-level generation and sentence-level rewards. Experimental results demonstrate that HPRO significantly enhances the emotional expressiveness of synthesized speech while preserving linguistic clarity and semantic integrity.

0 citationsRead paper

Divergence-Augmented Policy Optimization

Jan 25, 2025Neural Information Processing Systems

To address the instability and premature convergence caused by reusing offline data in deep reinforcement learning, this paper proposes a Bregman divergence constraint mechanism grounded in state distribution. Differing from conventional approaches that define Bregman divergence over action probability spaces, our method is the first to formulate it over the space of state distributions induced by policies, thereby establishing a divergence-augmented policy optimization framework. By explicitly constraining the magnitude of policy updates’ impact on the induced state distribution, the approach ensures both safety and efficacy in offline data reuse. Evaluated on the Atari benchmark under data-scarce settings, our method significantly improves training stability and convergence speed, while achieving superior sample efficiency and policy robustness compared to mainstream algorithms including PPO and SAC. These results empirically validate the effectiveness and practicality of regularization at the state-distribution level.

13 citations1 influentialRead paper