🤖 AI Summary
This study addresses the lack of bidirectional affective dynamic perception in embodied empathetic dialogue systems by proposing AffectLoop, a novel multimodal framework. Introducing a pioneering speaker-listener affective dynamics model, it integrates real-time emotional streams from both parties as conditional inputs for large language models, enabling closed-loop coordination between verbal and embodied behaviors. Within-subject experiments on the Misty II platform demonstrate that AffectLoop significantly outperforms baselines in empathetic response quality, user satisfaction, and emotional distress recovery. These findings validate the efficacy of the bidirectional affective feedback mechanism in enhancing alignment during human-robot empathetic interaction.
📝 Abstract
Empathetic social robots should respond not only to what users say, but also to how their emotions dynamically evolve during interaction. However, existing empathetic dialogue systems are often text-centered and primarily model empathy as a one-way mapping from the user's emotion to the system response, limiting their ability to capture embodied speaker--listener affective exchange. We present AffectLoop, a multimodal speaker-listener emotion-dynamics-aware spoken dialogue system implemented on the Misty II robot. The system tracks the speaker's verbal and facial affective dynamics, estimates the robot listener's own verbal and behavioral affective state, and conditions LLM-based response generation on both affective streams. The robot then generates a short spoken empathetic response together with emotionally congruent embodied behavior, forming a closed speaker--listener affective loop. We evaluate the system in a pilot within-subject study with five participants, comparing it with an otherwise identical utterance-conditioned baseline that omits the speaker- and listener-affective-state inputs. The proposed system received higher overall impression ratings, especially for empathetic response and user satisfaction. Post-hoc log analysis further showed higher speaker-listener affective alignment and stronger valence-based distress recovery. These preliminary results suggest that explicitly modeling both speaker emotional dynamics and listener affective state can improve embodied empathetic interaction.