E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing benchmarks for affective understanding predominantly evaluate expressed or induced emotions in isolation, lacking both fine-grained detail and comprehensiveness. This work proposes the first video-based affective understanding benchmark that jointly models induced and expressed emotions, comprising 12,314 question-answer pairs. It introduces three core tasks—emotion perception, open-vocabulary recognition, and Valence-Arousal-Dominance (VAD) dimensional assessment—to holistically evaluate multimodal large language models. The study innovatively employs a Bayesian pairwise alignment mechanism for efficient continuous annotation and designs E³mo-Score, a training-free multi-model committee scorer enabling fine-grained evaluation. Experimental results reveal significant performance disparities between induced and expressed emotion tasks and highlight persistent deficiencies in fine-grained recognition and dimensional modeling among current models.
📝 Abstract
Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interactions. However, existing benchmarks typically examine expressed and evoked emotions in isolation or are constrained to coarse-grained and incomplete affective characterizations. To bridge this gap, we introduce E$^3$mo-Bench, a scalable benchmark comprising $12{,}314$ question-answer pairs across $2{,}524$ videos with predefined affective perspectives. It evaluates evoked and expressed emotion understanding via $3$ complementary tasks: emotion perception, open-vocabulary recognition, and valence-arousal-dominance (VAD) assessment. To efficiently scale reliable continuous annotations, we propose Bayesian Pairwise Alignment, which aggregates sparse, low-burden pairwise judgments into anchor-referenced VAD estimates. Furthermore, we develop E$^3$mo-Score, a training-free agent that aggregates complementary judgments from a five-model committee to improve VAD estimation. Extensive experiments validate the effectiveness of our framework and expose a pronounced performance skew between evoked and expressed emotion paradigms. These findings, coupled with MLLMs' persistent deficits in fine-grained recognition and dimensional assessment, chart a clear course for advancing multimodal emotional intelligence.
Problem

Research questions and friction points this paper is trying to address.

multimodal emotion understanding
expressed emotion
evoked emotion
emotion benchmark
affective characterization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bayesian Pairwise Alignment
E³mo-Bench
multimodal emotion understanding
VAD estimation
training-free agent