Project Riley: Multimodal Multi-Agent LLM Collaboration with Emotional Reasoning and Voting

📅 2025-05-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations in multimodal perception, emotion modeling, and response coherence in affective conversational systems. We propose a personified multi-agent architecture inspired by *Inside Out*, comprising five emotion-specific agents—Joy, Sadness, Fear, Anger, and Disgust. The framework integrates multimodal large language models (text + vision), multi-turn negotiation, self-reflective iterative optimization, majority-voting fusion, and RAG-enhanced retrieval to enable emotion-driven dynamic perspective integration and logically coherent response generation. Key contributions include: (1) the first emotion-role-based multi-agent collaborative reasoning framework; (2) support for offline, lightweight deployment; and (3) an emergency-specialized variant, Armando, featuring cumulative context tracking and emotion-calibrated factual retrieval. User evaluations demonstrate statistically significant improvements over baselines in emotional appropriateness, expressive clarity, and naturalness, while maintaining real-time responsiveness.

Technology Category

Application Category

📝 Abstract
This paper presents Project Riley, a novel multimodal and multi-model conversational AI architecture oriented towards the simulation of reasoning influenced by emotional states. Drawing inspiration from Pixar's Inside Out, the system comprises five distinct emotional agents - Joy, Sadness, Fear, Anger, and Disgust - that engage in structured multi-round dialogues to generate, criticise, and iteratively refine responses. A final reasoning mechanism synthesises the contributions of these agents into a coherent output that either reflects the dominant emotion or integrates multiple perspectives. The architecture incorporates both textual and visual large language models (LLMs), alongside advanced reasoning and self-refinement processes. A functional prototype was deployed locally in an offline environment, optimised for emotional expressiveness and computational efficiency. From this initial prototype, another one emerged, called Armando, which was developed for use in emergency contexts, delivering emotionally calibrated and factually accurate information through the integration of Retrieval-Augmented Generation (RAG) and cumulative context tracking. The Project Riley prototype was evaluated through user testing, in which participants interacted with the chatbot and completed a structured questionnaire assessing three dimensions: Emotional Appropriateness, Clarity and Utility, and Naturalness and Human-likeness. The results indicate strong performance in structured scenarios, particularly with respect to emotional alignment and communicative clarity.
Problem

Research questions and friction points this paper is trying to address.

Simulating emotional reasoning in AI through multi-agent dialogues
Integrating multimodal LLMs for emotionally coherent outputs
Developing emergency AI with emotional calibration and factual accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal multi-agent LLM collaboration system
Emotional reasoning with voting mechanism
Integration of RAG and context tracking
💼 Related Jobs
No related jobs found.
A
Ana Rita Ortigoso
Computer Science and Communication Research Centre, Polytechnic University of Leiria
G
Gabriel Vieira
Computer Science and Communication Research Centre, Polytechnic University of Leiria
D
Daniel Fuentes
Computer Science and Communication Research Centre, Polytechnic University of Leiria
L
Luis Frazao
Computer Science and Communication Research Centre, Polytechnic University of Leiria
N
Nuno Costa
Computer Science and Communication Research Centre, Polytechnic University of Leiria
A
António Pereira
Computer Science and Communication Research Centre, Polytechnic University of Leiria