Multimodal Temporal Modeling for Continuous Group Emotion Recognition in Multi-party Dialogues

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对多参与者对话中群体情绪连续识别的问题,提出了一种结合音频和视频信息的多模态时序框架,并引入了混合状态来捕捉参与者间的情绪差异。
📝 Abstract
To realize natural behavior in dialogue agents in multi-party dialogue scenarios, it is important to understand group emotion such as valence and arousal as a whole. Most prior work addressed this task at the utterance level or using a coarse-grained time window, which is not sufficient to capture emotional dynamics. In this study, we formulate continuous recognition of the Group Emotion at a one-second resolution. Moreover, we also introduce the Mixed state, which captures the emotional divergence among participants in the group. We constructed a dataset with frame-level soft labels based on the TEIDAN corpus and propose a multimodal temporal framework that integrates audio and video information using a sliding-window context. Experimental results demonstrate that the temporal Transformer outperforms simple baselines and shows stronger temporal agreement with the ground-truth labels than the LLM-based model. The effect of context length is limited, whereas audio-visual input outperforms either unimodal input on the continuous-label metrics. Additionally, our analysis shows larger Group Emotion recognition errors in intervals with high Mixed values, exposing emotional divergence as a key challenge for group emotion recognition.
Problem

Research questions and friction points this paper is trying to address.

Group Emotion
Continuous Recognition
Multimodal Temporal Modeling
Emotional Dynamics
Mixed State
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Temporal Modeling
Continuous Group Emotion Recognition
Mixed State
Temporal Transformer
🔎 Similar Papers
No similar papers found.