Learning Discriminative Features from Spectrograms Using Center Loss for Speech Emotion Recognition
To address the challenges of ambiguous emotional representation and weak feature discriminability in speech emotion recognition (SER), this paper introduces center loss—a metric learning technique—into SER for the first time, proposing a joint optimization framework combining softmax cross-entropy loss and center loss. The method simultaneously enhances inter-class separability and intra-class compactness on variable-length Mel-spectrograms and STFT spectrograms. Leveraging a deep convolutional neural network, it directly learns highly discriminative emotional features from raw spectrograms without handcrafted features. Experiments on standard benchmark datasets demonstrate absolute improvements of 3.2% in unweighted accuracy and 4.1% in weighted accuracy over the softmax-only baseline. This work establishes a novel paradigm for emotion feature learning in SER and empirically validates the effectiveness of metric learning for improving discriminative capability in speech-based affective computing.