Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出DSSM-CRF模型,通过双向状态空间模型和动态条件随机场处理对话中的跨说话人影响及说话人内情绪演变问题,提高语音情感识别效果。
📝 Abstract
Conversational speech emotion recognition must reconcile acoustic evidence across temporal scales with two interaction processes: cross-speaker contextual influence and within-speaker emotion evolution. We propose DSSM-CRF, an audio-only architecture that explicitly separates these processes. Bidirectional state-space models encode fused self-supervised speech representations at frame and dialogue scales, so each utterance representation captures local prosody and context from all speakers. The decoder then orders each speaker's utterances into an independent dynamic conditional random field chain. Consecutive utterances in a speaker's chain form a transition pair whose score combines a corpus-level transition matrix with a residual predicted from the two contextualized utterances. An auxiliary objective supervises whether each pair changes emotion but does not participate in Viterbi inference. Thus, interlocutor turns affect contextual emotion scores without being treated as transitions in another speaker's emotion trajectory. DSSM-CRF achieves 75.81% UA and 74.90% WA on IEMOCAP, and 54.72% WA and 49.31% WF1 on MELD. Matched controls demonstrate complementary gains from speaker-wise factorization and CRF modeling.
Problem

Research questions and friction points this paper is trying to address.

Speech Emotion Recognition
Conversational Speech
State-Space Modeling
Cross-Speaker Contextual Influence
Within-Speaker Emotion Evolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-Scale State-Space Modeling
Speaker-Wise Dynamic CRF
Speech Emotion Recognition
Bidirectional state-space models
Dynamic conditional random field
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Guan-Hua Wen
Department of Computer Science and Information Engineering, National Taiwan University of Science and Technology, Taipei, Taiwan
Kuan-Yu Chen
Kuan-Yu Chen
National Taiwan University of Science and Technology
Language ModelingSpeech RecognitionInformation RetrievalSummarizationNature Language Processing
H
Hou-Chiang Tseng
Graduate Institute of Digital Learning and Education, National Taiwan University of Science and Technology, Taipei, Taiwan