Multimodal Emotion Recognition in Conversations via Class-Wise Adaptive Modality Fusion and Affective Geometry

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过结合面部几何特征、类别自适应模态融合和情感几何先验的方法,提高了对话中多模态情感识别的准确性。
📝 Abstract
Emotion Recognition in Conversations (ERC) requires integrating heterogeneous textual, audio, and visual cues while accounting for conversational context and emotional dynamics. We extend the Self-Distillation Transformer architecture for ERC with appearance+geometry visual representations, class-wise adaptive modality fusion, and a valence-arousal prior for affective transitions. On the MELD and IEMOCAP datasets, geometry-enhanced visual representations improve weighted F1 by 0.27 and 4.36 points over appearance-only features, respectively, while class-wise adaptive fusion provides further gains of 0.17 and 0.25 points over the original softmax gate. The valence-arousal prior yields targeted improvements of 0.30 and 0.74 accuracy points on emotionally shifted utterances while preserving performance on stable turns. These results indicate that structured facial cues, emotion-dependent modality weighting, and affective geometry provide complementary benefits for multimodal ERC.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Emotion Recognition
Conversational Context
Emotional Dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

class-wise adaptive modality fusion
affective geometry
valence-arousal prior
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
O
Oriol Marín
Eurecat, Centre Tecnològic de Catalunya, Barcelona, Spain
R
Roger Marí
Eurecat, Centre Tecnològic de Catalunya, Barcelona, Spain
Gloria Haro
Gloria Haro
Universitat Pompeu Fabra
Computer VisionAudio-visual ProcessingDeep Learning
R
Rafael Redondo
Eurecat, Centre Tecnològic de Catalunya, Barcelona, Spain