🤖 AI Summary
Existing benchmarks for affective understanding predominantly evaluate expressed or induced emotions in isolation, lacking both fine-grained detail and comprehensiveness. This work proposes the first video-based affective understanding benchmark that jointly models induced and expressed emotions, comprising 12,314 question-answer pairs. It introduces three core tasks—emotion perception, open-vocabulary recognition, and Valence-Arousal-Dominance (VAD) dimensional assessment—to holistically evaluate multimodal large language models. The study innovatively employs a Bayesian pairwise alignment mechanism for efficient continuous annotation and designs E³mo-Score, a training-free multi-model committee scorer enabling fine-grained evaluation. Experimental results reveal significant performance disparities between induced and expressed emotion tasks and highlight persistent deficiencies in fine-grained recognition and dimensional modeling among current models.
📝 Abstract
Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interactions. However, existing benchmarks typically examine expressed and evoked emotions in isolation or are constrained to coarse-grained and incomplete affective characterizations. To bridge this gap, we introduce E$^3$mo-Bench, a scalable benchmark comprising $12{,}314$ question-answer pairs across $2{,}524$ videos with predefined affective perspectives. It evaluates evoked and expressed emotion understanding via $3$ complementary tasks: emotion perception, open-vocabulary recognition, and valence-arousal-dominance (VAD) assessment. To efficiently scale reliable continuous annotations, we propose Bayesian Pairwise Alignment, which aggregates sparse, low-burden pairwise judgments into anchor-referenced VAD estimates. Furthermore, we develop E$^3$mo-Score, a training-free agent that aggregates complementary judgments from a five-model committee to improve VAD estimation. Extensive experiments validate the effectiveness of our framework and expose a pronounced performance skew between evoked and expressed emotion paradigms. These findings, coupled with MLLMs' persistent deficits in fine-grained recognition and dimensional assessment, chart a clear course for advancing multimodal emotional intelligence.