EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing speech-language models in emotional intelligence evaluation, which often relies on shallow paralinguistic features without grounding in cognitive theory. Drawing upon the four-branch model of emotional intelligence, the study introduces EmoSBench—the first theory-driven benchmark—and proposes EmoS, a novel model optimized through supervised fine-tuning (SFT) and Grouped Relative Policy Optimization (GRPO). EmoS incorporates a joint reward mechanism combining Steep Exponential Accuracy Reward (SEAR) and Reasoning Fidelity Reward (RFR), trained on EmoDialogue, a newly curated fine-grained bilingual conversational dataset. Experimental results demonstrate that EmoS achieves 83.8% accuracy on EmoSBench, approaching human-level performance, and exhibits strong generalization capabilities in real-world, unconstrained spoken interactions.
📝 Abstract
Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a systematic, theory-driven cognitive framework. We introduce EmoSBench, the first comprehensive EI evaluation benchmark for SLMs constructed upon the four-branch theoretical model, covering Perceiving, Understanding, Using, and Managing Emotion across ten sub-tasks. Preliminary assessments on EmoSBench reveal a substantial gap: even leading proprietary models like GPT-4o-Audio achieve only 52.6%, significantly trailing human baselines. To bridge this gap, we develop EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To facilitate its effective training, we curate EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations. Concurrently, we introduce a reward mechanism integrating a Steep Exponential Accuracy Reward (SEAR) and a Rationale Fidelity Reward (RFR) to enforce precise ordinal scoring and valid reasoning. Experiments demonstrate that EmoS reaches 83.8% accuracy, approaching human-level performance. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust real-world generalization, establishing a foundational framework for advancing emotionally intelligent dialogue systems.
Problem

Research questions and friction points this paper is trying to address.

Emotional Intelligence
Spoken Language Models
Evaluation Benchmark
Theory-Grounded Framework
Paralinguistic Perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

Emotional Intelligence
Spoken Language Models
Theory-Grounded Evaluation
Group Relative Policy Optimization
Reward Mechanism
🔎 Similar Papers
No similar papers found.
J
Junyu Wang
Tianjin Key Laboratory of Cognitive Computing and Application, Tianjin University, Tianjin, China
S
Siyuan Zhang
Tianjin Key Laboratory of Cognitive Computing and Application, Tianjin University, Tianjin, China
P
Peiyuan Jiang
Tianjin Key Laboratory of Cognitive Computing and Application, Tianjin University, Tianjin, China
J
Jian Zong
Tianjin Key Laboratory of Cognitive Computing and Application, Tianjin University, Tianjin, China
Jingyu Zhang
Jingyu Zhang
WNLO Huazhong University of Science and Technology
optical
Tianrui Wang
Tianrui Wang
Tianjin University
Speech Signal Processing
Y
Yuqin Lin
Fuzhou University, Fuzhou, China
Z
Zhenghui Chen
Fuzhou University, Fuzhou, China
S
Shuqing Xie
Fuzhou University, Fuzhou, China
Ziyang Ma
Ziyang Ma
Shanghai Jiao Tong University
Speech and Language ProcessingTextless NLPSelf-supervised LearningMultimedia
Meng Ge
Meng Ge
Tianjin University; CUHK-Shenzhen; National University of Singapore
Xiaobao Wang
Xiaobao Wang
天津大学 Associate Professor
人工智能,大模型生成安全,图机器学习
Longbiao Wang
Longbiao Wang
Professor, Tianjin University
Speech ProcessingSpeech recognitionspeaker recognitionacoustic signal processingspeech enhancement
Jianwu Dang
Jianwu Dang
JAIST, Japan / Tianjin Univ., China
Speech Sciencespeech productionEEGdisorder speech