ESCRAG-R1: Retrieval-Augmented Reinforcement Learning for Emotional Support Conversation

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决情感支持对话中专业性和同理心难以兼顾的问题,提出ESCRAG-R1框架,结合检索增强的强化学习和高质量评估数据集ESC-Preference,实现自然融合。
📝 Abstract
Emotional Support Conversation (ESC) systems aim to provide holistic support by balancing professional therapeutic competence with natural empathy. However, existing methods struggle to simultaneously achieve structured, stage-aware reasoning and seamless empathy-expertise alignment, often resulting in an artificial splicing of clinical strategies and generic reassurance. To overcome these limitations, we propose ESCRAG-R1, a unified framework that integrates retrieval-based psychological guidance into Group Relative Policy Optimization (GRPO). By incorporating retrieval into the reinforcement learning loop, ESCRAG-R1 transforms external knowledge into a robust learning signal that stimulates explicit internal reasoning prior to generation and fundamentally reshapes the model's internal policy. To provide the reliable supervision required for this optimization, we construct ESC-Preference, a high-quality dataset based on a Client--Counselor--Judge evaluation framework that delivers precise, empathy-aware reward signals. Extensive experiments demonstrate that ESCRAG-R1 significantly outperforms existing baselines by mitigating superficial splicing and realizing a natural integration of professional guidance and empathetic expression. Code and datasets are released at https://github.com/Matcha-Liu/ESCRAG-R1.
Problem

Research questions and friction points this paper is trying to address.

Emotional Support Conversation
reinforcement learning
empathy-expertise alignment
structured reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Augmented Reinforcement Learning
Emotional Support Conversation
Group Relative Policy Optimization (GRPO)
ESC-Preference Dataset
🔎 Similar Papers
2024-02-20Annual Meeting of the Association for Computational LinguisticsCitations: 17
W
Weichu Liu
Beijing Institute of Technology
Yuxuan Hu
Yuxuan Hu
The Chinese University of Hong Kong
multimodalLLM
Y
Yirong Sun
Shenzhen University of Advanced Technology
N
Ningning Mao
Beijing Normal University
Z
Ziyun Zhang
Beijing Institute of Technology
J
Jian Chen
The University of Hong Kong
M
Mingyang Xu
Shenzhen MSU-BIT University
Q
Qishan Zhong
Shenzhen MSU-BIT University
C
Chengming Li
Shenzhen MSU-BIT University