🤖 AI Summary
This work addresses the limitation of existing speech-language models in emotional intelligence evaluation, which often relies on shallow paralinguistic features without grounding in cognitive theory. Drawing upon the four-branch model of emotional intelligence, the study introduces EmoSBench—the first theory-driven benchmark—and proposes EmoS, a novel model optimized through supervised fine-tuning (SFT) and Grouped Relative Policy Optimization (GRPO). EmoS incorporates a joint reward mechanism combining Steep Exponential Accuracy Reward (SEAR) and Reasoning Fidelity Reward (RFR), trained on EmoDialogue, a newly curated fine-grained bilingual conversational dataset. Experimental results demonstrate that EmoS achieves 83.8% accuracy on EmoSBench, approaching human-level performance, and exhibits strong generalization capabilities in real-world, unconstrained spoken interactions.
📝 Abstract
Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a systematic, theory-driven cognitive framework. We introduce EmoSBench, the first comprehensive EI evaluation benchmark for SLMs constructed upon the four-branch theoretical model, covering Perceiving, Understanding, Using, and Managing Emotion across ten sub-tasks. Preliminary assessments on EmoSBench reveal a substantial gap: even leading proprietary models like GPT-4o-Audio achieve only 52.6%, significantly trailing human baselines. To bridge this gap, we develop EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To facilitate its effective training, we curate EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations. Concurrently, we introduce a reward mechanism integrating a Steep Exponential Accuracy Reward (SEAR) and a Rationale Fidelity Reward (RFR) to enforce precise ordinal scoring and valid reasoning. Experiments demonstrate that EmoS reaches 83.8% accuracy, approaching human-level performance. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust real-world generalization, establishing a foundational framework for advancing emotionally intelligent dialogue systems.