Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey

📅 2025-09-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
A systematic survey of multimodal large language models (MLLMs) for multimodal sentiment recognition and reasoning remains absent. Method: This work presents the first comprehensive review of this interdisciplinary domain, covering mainstream MLLM architectures, multimodal sentiment datasets, benchmark performance, and technical pathways for cross-modal fusion and semantic reasoning. Through rigorous literature screening and horizontal comparative analysis, we identify core challenges—including insufficient model robustness, difficulties in cross-modal alignment, and weak fine-grained sentiment modeling. Contribution/Results: We propose promising future directions: enhanced interpretability, dynamic modality weighting, and causal sentiment reasoning. Concurrently, we release an open-source resource repository—including curated models, datasets, and evaluation code—to provide researchers with an authoritative survey, practical guidelines, and reproducible benchmarks—thereby advancing MLLM-driven affective intelligence.

Technology Category

Application Category

📝 Abstract
In recent years, large language models (LLMs) have driven major advances in language understanding, marking a significant step toward artificial general intelligence (AGI). With increasing demands for higher-level semantics and cross-modal fusion, multimodal large language models (MLLMs) have emerged, integrating diverse information sources (e.g., text, vision, and audio) to enhance modeling and reasoning in complex scenarios. In AI for Science, multimodal emotion recognition and reasoning has become a rapidly growing frontier. While LLMs and MLLMs have achieved notable progress in this area, the field still lacks a systematic review that consolidates recent developments. To address this gap, this paper provides a comprehensive survey of LLMs and MLLMs for emotion recognition and reasoning, covering model architectures, datasets, and performance benchmarks. We further highlight key challenges and outline future research directions, aiming to offer researchers both an authoritative reference and practical insights for advancing this domain. To the best of our knowledge, this paper is the first attempt to comprehensively survey the intersection of MLLMs with multimodal emotion recognition and reasoning. The summary of existing methods mentioned is in our Github: href{https://github.com/yuntaoshou/Awesome-Emotion-Reasoning}{https://github.com/yuntaoshou/Awesome-Emotion-Reasoning}.
Problem

Research questions and friction points this paper is trying to address.

Surveying multimodal emotion recognition using large language models
Addressing lack of systematic review in multimodal emotion reasoning
Providing comprehensive analysis of architectures, datasets and benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrating text, vision, and audio data
Surveying model architectures and performance benchmarks
Providing comprehensive review of emotion recognition methods
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.