Memo2496: Expert-Annotated Dataset and Dual-View Adaptive Framework for Music Emotion Recognition

📅 2025-12-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Music emotion recognition (MER) faces two key challenges: scarcity of high-quality annotated data and cross-track feature drift. To address these, we introduce Memo2496—the first large-scale expert-annotated instrumental emotion dataset comprising 2,496 tracks with continuous valence-arousal labels—and propose DAMER, a dual-view adaptive framework. Its core contributions are: (1) a dual-stream attention mechanism fusing Mel-spectrogram and cochleagram representations; (2) progressive confidence-based pseudo-labeling, integrating extreme-emotion calibration and consistency filtering with a threshold of 0.25; and (3) style-anchored contrastive memory learning. DAMER further incorporates curriculum-learning-inspired temperature scheduling and Jensen–Shannon divergence to quantify inter-view consistency. Evaluated on Memo2496, 1000songs, and PMEmo, DAMER achieves state-of-the-art arousal classification accuracy—improving by 3.43%, 2.25%, and 0.17%, respectively. Both the Memo2496 dataset and source code are fully open-sourced.

Technology Category

Application Category

📝 Abstract
Music Emotion Recogniser (MER) research faces challenges due to limited high-quality annotated datasets and difficulties in addressing cross-track feature drift. This work presents two primary contributions to address these issues. Memo2496, a large-scale dataset, offers 2496 instrumental music tracks with continuous valence arousal labels, annotated by 30 certified music specialists. Annotation quality is ensured through calibration with extreme emotion exemplars and a consistency threshold of 0.25, measured by Euclidean distance in the valence arousal space. Furthermore, the Dual-view Adaptive Music Emotion Recogniser (DAMER) is introduced. DAMER integrates three synergistic modules: Dual Stream Attention Fusion (DSAF) facilitates token-level bidirectional interaction between Mel spectrograms and cochleagrams via cross attention mechanisms; Progressive Confidence Labelling (PCL) generates reliable pseudo labels employing curriculum-based temperature scheduling and consistency quantification using Jensen Shannon divergence; and Style Anchored Memory Learning (SAML) maintains a contrastive memory queue to mitigate cross-track feature drift. Extensive experiments on the Memo2496, 1000songs, and PMEmo datasets demonstrate DAMER's state-of-the-art performance, improving arousal dimension accuracy by 3.43%, 2.25%, and 0.17%, respectively. Ablation studies and visualisation analyses validate each module's contribution. Both the dataset and source code are publicly available.
Problem

Research questions and friction points this paper is trying to address.

Addresses limited high-quality annotated datasets for music emotion recognition
Mitigates cross-track feature drift in music emotion recognition models
Improves accuracy in continuous valence-arousal emotion prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-view attention fusion for spectrogram-cochleagram interaction
Progressive confidence labeling with curriculum temperature scheduling
Style-anchored memory learning to mitigate cross-track feature drift
💼 Related Jobs
No related jobs found.
Q
Qilin Li
Guangdong Provincial Key Laboratory of AI Large Model and Intelligent Cognition, the School of Computer Science and Engineering, South China University of Technology, Guangzhou 510006, China, and with Research Centre for AI Large Models and Intelligent Cognition, Pazhou Lab, Guangzhou 510335, China, and also with Engineering Research Centre of the Ministry of Education on Health Intelligent Perception and Paralleled Digital-Human, South China University of Technology, Guangzhou 510006, China
C
C. L. Philip Chen
Guangdong Provincial Key Laboratory of AI Large Model and Intelligent Cognition, the School of Computer Science and Engineering, South China University of Technology, Guangzhou 510006, China, and with Research Centre for AI Large Models and Intelligent Cognition, Pazhou Lab, Guangzhou 510335, China, and also with Engineering Research Centre of the Ministry of Education on Health Intelligent Perception and Paralleled Digital-Human, South China University of Technology, Guangzhou 510006, China
T
Tong Zhang
Guangdong Provincial Key Laboratory of AI Large Model and Intelligent Cognition, the School of Computer Science and Engineering, South China University of Technology, Guangzhou 510006, China, and with Research Centre for AI Large Models and Intelligent Cognition, Pazhou Lab, Guangzhou 510335, China, and also with Engineering Research Centre of the Ministry of Education on Health Intelligent Perception and Paralleled Digital-Human, South China University of Technology, Guangzhou 510006, China