Institution profile

Chinese University of Hong Kong

Academic institutionasia · hk
Official website
Research library3,784linked papers
Opportunities0open roles
Selected work

Representative Papers

Learning Discriminative Features from Spectrograms Using Center Loss for Speech Emotion Recognition

May 01, 2019IEEE International Conference on Acoustics, Speech, and Signal Processing

To address the challenges of ambiguous emotional representation and weak feature discriminability in speech emotion recognition (SER), this paper introduces center loss—a metric learning technique—into SER for the first time, proposing a joint optimization framework combining softmax cross-entropy loss and center loss. The method simultaneously enhances inter-class separability and intra-class compactness on variable-length Mel-spectrograms and STFT spectrograms. Leveraging a deep convolutional neural network, it directly learns highly discriminative emotional features from raw spectrograms without handcrafted features. Experiments on standard benchmark datasets demonstrate absolute improvements of 3.2% in unweighted accuracy and 4.1% in weighted accuracy over the softmax-only baseline. This work establishes a novel paradigm for emotion feature learning in SER and empirically validates the effectiveness of metric learning for improving discriminative capability in speech-based affective computing.

48 citations6 influentialRead paper

Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-Trained BERT

Sep 15, 2019Interspeech

This paper addresses the polyphonic character disambiguation challenge in Chinese end-to-end speech synthesis. We propose the first end-to-end grapheme-to-phoneme (G2P) framework for Chinese leveraging pretrained BERT: it takes raw character sequences as input—requiring neither manual word segmentation nor phoneme pre-annotation—and employs BERT to encode contextual semantics, jointly with FC, LSTM, and Transformer classifiers to predict polyphonic pronunciations. Our key contribution is the first application of BERT to Chinese polyphonic disambiguation, enabling semantic-driven, end-to-end modeling. Experiments demonstrate that BERT substantially improves disambiguation accuracy, outperforming an LSTM baseline on a standard test set. Furthermore, we empirically reveal that context window length critically influences disambiguation performance.

24 citationsRead paper

A Survey on Vision-Language-Action Models for Embodied AI

May 23, 2024arXiv.org

This paper addresses the core challenge of how Vision-Language-Action (VLA) models support language-conditioned robotic tasks in embodied intelligence. Methodologically, it introduces the first systematic, panoramic survey framework, proposing a three-dimensional taxonomy—“Component Design–Low-level Action Policies–High-level Task Planning”—that unifies VLA modeling, embodied control, task decomposition, simulation integration, and cross-benchmark evaluation. Key contributions include: (1) the first explicit characterization of three principal VLA technical paradigms; (2) a comprehensive survey of multimodal datasets, embodied simulation platforms, and standardized evaluation benchmarks; and (3) a structured knowledge graph that identifies critical open challenges—including scalable architecture design, world model integration, and real-world deployment—and outlines promising future research directions.

18 citations1 influentialRead paper

Divergence-Augmented Policy Optimization

Jan 25, 2025Neural Information Processing Systems

To address the instability and premature convergence caused by reusing offline data in deep reinforcement learning, this paper proposes a Bregman divergence constraint mechanism grounded in state distribution. Differing from conventional approaches that define Bregman divergence over action probability spaces, our method is the first to formulate it over the space of state distributions induced by policies, thereby establishing a divergence-augmented policy optimization framework. By explicitly constraining the magnitude of policy updates’ impact on the induced state distribution, the approach ensures both safety and efficacy in offline data reuse. Evaluated on the Atari benchmark under data-scarce settings, our method significantly improves training stability and convergence speed, while achieving superior sample efficiency and policy robustness compared to mainstream algorithms including PPO and SAC. These results empirically validate the effectiveness and practicality of regularization at the state-distribution level.

13 citations1 influentialRead paper

Data-Driven Merton's Strategies via Policy Randomization

Dec 19, 2023

This paper addresses the Merton expected utility maximization problem in an incomplete market with fully unknown dynamics: the market comprises a stock and latent state-factor processes, while investors observe only prices and instantaneous volatility—rendering factor dynamics and market parameters unidentifiable. We propose a novel continuous-time reinforcement learning (RL) framework that, for the first time, employs Gaussian policy randomization as an analytical tool—not merely for exploration—and rigorously prove that the randomized policy’s mean coincides with the optimal control of the original problem, thereby bridging RL and classical portfolio theory. We design online and offline actor-critic algorithms integrating policy improvement theorems with randomized policy analysis. Under stochastic volatility, our method substantially outperforms conventional parametric interpolation approaches. Both simulation studies and empirical tests confirm its robustness and capacity to generate alpha.

10 citations1 influentialRead paper
Recent publications

Latest Papers