Emotion as a Distribution: Joint Valence-Arousal Probability Learning for Speaker-Independent Multimodal Emotion Recognition

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Human emotion is graded and frequently mixed, yet most multimodal recognizers collapse it onto a single hard label. We argue the recognizer should instead expose a distribution over affective space. Our text+speech system, alongside its categorical decision, emits a $9\times9$ probability matrix over the Valence-Arousal plane, trained with a two-dimensional Gaussian soft target under a Kullback-Leibler/cross-entropy objective, aimed at counseling support. Evaluation is strict: speaker-independent 5-fold leave-one-session-out IEMOCAP with rotating-session inner validation, headline metrics only on the held-out session. Within one fixed encoder-fusion-head pipeline we compare Transformer and state-space (Mamba-1/2/3) backbones at matched depth and width, at two operating points ($T\approx550$, $T\approx2750$). The featured dual-head system reaches 73.0% $\pm$ 0.3 unweighted accuracy over three seeds (separate rerun: 72.1%), exceeding the Transformer fusion baseline by 3.0 UA points (95% session-bootstrap CI [1.0,4.7]; significant under paired t-test and session-level bootstrap), with no latency or memory advantage at these lengths; swapping the ~1M trainable front-end for frozen WavLM-Large features (learnable layer weights) lifts the same architecture to 76.6% $\pm$ 1.3. Pre-specified controls scope the claims honestly: simpler valence-arousal auxiliaries reproduce the classification lift within noise, and a dedicated regression head tracks the continuous ratings slightly better, so the head's specific value is the normalized affect distribution itself. That distribution recovers the circumplex: its center of mass tracks valence and arousal (CCC 0.66/0.66; predominantly between-class structure, weaker within-class tracking), and its entropy is weakly but consistently linked to categorical rater ambiguity, not dimensional spread.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Emotion Recognition
Valence-Arousal Distribution
Speaker-Independent
Innovation

Methods, ideas, or system contributions that make the work stand out.

Valence-Arousal Probability Learning
Multimodal Emotion Recognition
Two-dimensional Gaussian Soft Target
Kullback-Leibler/Cross-Entropy Objective
Speaker-Independent Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tingyi Lin
Department of Electrical Engineering, National Changhua University of Education, Changhua, Taiwan
W
Wen-Ren Yang
Department of Electrical Engineering, National Changhua University of Education, Changhua, Taiwan
K
Kuanwei Chen
Department of Computer Science and Information Engineering, National Central University, Taoyuan, Taiwan