Silicon Minds versus Human Hearts: The Wisdom of Crowds Beats the Wisdom of AI in Emotion Recognition

📅 2025-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether multimodal large language models (MLLMs) can match or surpass human experts in emotion recognition, and evaluates the performance gains from human-AI collaboration and collective intelligence. Using the standardized Reading the Mind in the Eyes Test (RMET) and its multi-ethnic extension (MRMET), we benchmark individual MLLMs, individual humans, human groups (aggregating independent judgments), and human-AI collaborative frameworks. Results show that individual MLLMs significantly outperform individual humans in accuracy; however, human group decisions consistently exceed all single-model baselines. A novel human-AI co-reasoning framework—integrating model-generated explanations with human group consensus—achieves the highest overall accuracy. This work is the first to systematically demonstrate the superiority of collective intelligence in affective understanding, introducing the “augmented intelligence” paradigm. It provides both theoretical foundations and practical design principles for developing trustworthy, human-aligned affective AI systems.

Technology Category

Application Category

📝 Abstract
The ability to discern subtle emotional cues is fundamental to human social intelligence. As artificial intelligence (AI) becomes increasingly common, AI's ability to recognize and respond to human emotions is crucial for effective human-AI interactions. In particular, whether such systems can match or surpass human experts remains to be seen. However, the emotional intelligence of AI, particularly multimodal large language models (MLLMs), remains largely unexplored. This study evaluates the emotion recognition abilities of MLLMs using the Reading the Mind in the Eyes Test (RMET) and its multiracial counterpart (MRMET), and compares their performance against human participants. Results show that, on average, MLLMs outperform humans in accurately identifying emotions across both tests. This trend persists even when comparing performance across low, medium, and expert-level performing groups. Yet when we aggregate independent human decisions to simulate collective intelligence, human groups significantly surpass the performance of aggregated MLLM predictions, highlighting the wisdom of the crowd. Moreover, a collaborative approach (augmented intelligence) that combines human and MLLM predictions achieves greater accuracy than either humans or MLLMs alone. These results suggest that while MLLMs exhibit strong emotion recognition at the individual level, the collective intelligence of humans and the synergistic potential of human-AI collaboration offer the most promising path toward effective emotional AI. We discuss the implications of these findings for the development of emotionally intelligent AI systems and future research directions.
Problem

Research questions and friction points this paper is trying to address.

Evaluates MLLMs' emotion recognition vs humans using RMET/MRMET tests
Compares individual and collective human performance with AI predictions
Explores augmented intelligence combining human and MLLM emotion detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

MLLMs outperform humans in emotion recognition
Human crowds surpass aggregated MLLM predictions
Human-AI collaboration achieves highest accuracy
🔎 Similar Papers
No similar papers found.