EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态对话数据集情感标注不足、情感多样性差和规模小的问题,通过设计新的数据收集框架创建了大规模音频视觉情感数据集EmotionDialogCN。
📝 Abstract
Face-to-face audiovisual interaction is central to human communication, conveying rich emotional and social cues. However, existing multimodal dialogue datasets remain limited by inadequate emotion annotations, poor emotional diversity, and small scale. We introduce EmotionDialogCN, a large-scale audiovisual-emotional dataset designed to capture authentic face-to-face communication. It contains 21,880 dialogue sessions performed by 119 professional actors across 20 everyday scenarios, covering 18 emotion categories with over 400 hours of recordings, the largest and most comprehensive dataset of its kind. A novel data collection framework minimizes equipment interference, enabling natural and nuanced emotional expressions. EmotionDialogCN achieves an emotion distribution deviation of 0.64 from real human emotion statistics (versus 5.65 for prior datasets) and consistent subject framing (52-59% frame occupancy). Together, these properties translate into stable unimodal and multimodal performance across acoustic, lexical, and visual modalities, with fusion results further underscoring strong multimodal alignment and cross-modal complementarity.
Problem

Research questions and friction points this paper is trying to address.

multimodal dialogue datasets
emotion annotations
emotional diversity
Innovation

Methods, ideas, or system contributions that make the work stand out.

large-scale dataset
multimodal emotional dialogue
natural expressions
emotion distribution accuracy
cross-modal complementarity
🔎 Similar Papers
No similar papers found.