Institution profile

Central South University of Forestry and Technology

Academic institutionasia · cn
Official website
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

Dual-branch Graph Domain Adaptation for Cross-scenario Multi-modal Emotion Recognition

Mar 27, 2026

This work addresses the limited generalization of multimodal conversational emotion recognition across diverse scenarios, where variations in speakers, topics, styles, and noise degrade performance. To tackle this challenge, the authors propose a dual-branch graph-based domain adaptation framework that jointly models domain adaptation and robustness to label noise for the first time. The method constructs an emotion interaction hypergraph and employs a dual-branch encoder to capture both local multi-way relationships and global dependencies. A domain adversarial discriminator is integrated to learn domain-invariant representations, while a regularization loss mitigates the adverse effects of label noise. Theoretical analysis yields a tighter generalization bound. Extensive experiments on IEMOCAP and MELD demonstrate that the proposed model significantly outperforms strong baselines, achieving superior cross-scenario emotion recognition and generalization capabilities.

0 citationsRead paper

Dynamic Fusion-Aware Graph Convolutional Neural Network for Multimodal Emotion Recognition in Conversations

Mar 21, 2026

Existing multimodal conversational emotion recognition methods typically employ fixed-parameter fusion of multimodal features, which struggles to accommodate the dynamic requirements of different emotion categories and thus limits recognition performance. To address this limitation, this work proposes a Dynamic Fusion-aware Graph Convolutional Network (DF-GCN), which introduces ordinary differential equations into the graph convolutional framework for the first time and designs a global information-guided dynamic prompting mechanism that adaptively adjusts fusion parameters according to the target emotion category. Extensive experiments on two public multimodal conversational datasets demonstrate that the proposed method significantly outperforms state-of-the-art models, validating the effectiveness of the dynamic fusion mechanism in enhancing both accuracy and generalization capability in emotion recognition.

0 citationsRead paper

Relational graph-driven differential denoising and diffusion attention fusion for multimodal conversation emotion recognition

Mar 21, 2026

This work addresses the challenge of multimodal conversational emotion recognition under noisy environmental and acquisition conditions, which often introduce noise into audiovisual features and cause imbalanced information quality across modalities, leading to distorted representations and biased weighting during fusion. To mitigate these issues, the authors propose a relation-aware denoising and diffusion attention fusion model. The approach employs a differential Transformer to suppress temporally irrelevant noise, constructs intra- and inter-modal relation subgraphs to capture affective dependencies, and introduces a text-guided cross-modal diffusion mechanism to achieve semantically aligned and robust fusion. Notably, this method explicitly models modality discrepancies and the dominant role of textual cues—departing from conventional implicit weighted fusion schemes—and demonstrates significant improvements in both accuracy and robustness for emotion recognition in noisy settings.

0 citationsRead paper

TimeGNN-Augmented Hybrid-Action MARL for Fine-Grained Task Partitioning and Energy-Aware Offloading in MEC

Jan 08, 2026arXiv.org

This work addresses the challenges of task scheduling and resource allocation in mobile edge computing, where limited resources, unstable power supply, and high system dynamics pose significant difficulties. To tackle these issues, the authors propose TG-DCMADDPG, a multi-agent reinforcement learning framework that integrates a Temporal Graph Neural Network (TimeGNN) to model multidimensional temporal states and introduces a discrete-continuous hybrid action space within a multi-agent deep deterministic policy gradient algorithm. This approach jointly optimizes fine-grained task partitioning, transmission power allocation, and scheduling priorities. Experimental results demonstrate that the proposed method significantly accelerates policy convergence, improves energy efficiency and latency performance, enhances task completion rates, and exhibits strong scalability and practical applicability.

0 citationsRead paper

Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey

Sep 29, 2025

A systematic survey of multimodal large language models (MLLMs) for multimodal sentiment recognition and reasoning remains absent. Method: This work presents the first comprehensive review of this interdisciplinary domain, covering mainstream MLLM architectures, multimodal sentiment datasets, benchmark performance, and technical pathways for cross-modal fusion and semantic reasoning. Through rigorous literature screening and horizontal comparative analysis, we identify core challenges—including insufficient model robustness, difficulties in cross-modal alignment, and weak fine-grained sentiment modeling. Contribution/Results: We propose promising future directions: enhanced interpretability, dynamic modality weighting, and causal sentiment reasoning. Concurrently, we release an open-source resource repository—including curated models, datasets, and evaluation code—to provide researchers with an authoritative survey, practical guidelines, and reproducible benchmarks—thereby advancing MLLM-driven affective intelligence.

0 citationsRead paper
Recent publications

Latest Papers

Dual-branch Graph Domain Adaptation for Cross-scenario Multi-modal Emotion Recognition

Mar 27, 2026

This work addresses the limited generalization of multimodal conversational emotion recognition across diverse scenarios, where variations in speakers, topics, styles, and noise degrade performance. To tackle this challenge, the authors propose a dual-branch graph-based domain adaptation framework that jointly models domain adaptation and robustness to label noise for the first time. The method constructs an emotion interaction hypergraph and employs a dual-branch encoder to capture both local multi-way relationships and global dependencies. A domain adversarial discriminator is integrated to learn domain-invariant representations, while a regularization loss mitigates the adverse effects of label noise. Theoretical analysis yields a tighter generalization bound. Extensive experiments on IEMOCAP and MELD demonstrate that the proposed model significantly outperforms strong baselines, achieving superior cross-scenario emotion recognition and generalization capabilities.

0 citationsRead paper

Dynamic Fusion-Aware Graph Convolutional Neural Network for Multimodal Emotion Recognition in Conversations

Mar 21, 2026

Existing multimodal conversational emotion recognition methods typically employ fixed-parameter fusion of multimodal features, which struggles to accommodate the dynamic requirements of different emotion categories and thus limits recognition performance. To address this limitation, this work proposes a Dynamic Fusion-aware Graph Convolutional Network (DF-GCN), which introduces ordinary differential equations into the graph convolutional framework for the first time and designs a global information-guided dynamic prompting mechanism that adaptively adjusts fusion parameters according to the target emotion category. Extensive experiments on two public multimodal conversational datasets demonstrate that the proposed method significantly outperforms state-of-the-art models, validating the effectiveness of the dynamic fusion mechanism in enhancing both accuracy and generalization capability in emotion recognition.

0 citationsRead paper

Relational graph-driven differential denoising and diffusion attention fusion for multimodal conversation emotion recognition

Mar 21, 2026

This work addresses the challenge of multimodal conversational emotion recognition under noisy environmental and acquisition conditions, which often introduce noise into audiovisual features and cause imbalanced information quality across modalities, leading to distorted representations and biased weighting during fusion. To mitigate these issues, the authors propose a relation-aware denoising and diffusion attention fusion model. The approach employs a differential Transformer to suppress temporally irrelevant noise, constructs intra- and inter-modal relation subgraphs to capture affective dependencies, and introduces a text-guided cross-modal diffusion mechanism to achieve semantically aligned and robust fusion. Notably, this method explicitly models modality discrepancies and the dominant role of textual cues—departing from conventional implicit weighted fusion schemes—and demonstrates significant improvements in both accuracy and robustness for emotion recognition in noisy settings.

0 citationsRead paper

TimeGNN-Augmented Hybrid-Action MARL for Fine-Grained Task Partitioning and Energy-Aware Offloading in MEC

Jan 08, 2026arXiv.org

This work addresses the challenges of task scheduling and resource allocation in mobile edge computing, where limited resources, unstable power supply, and high system dynamics pose significant difficulties. To tackle these issues, the authors propose TG-DCMADDPG, a multi-agent reinforcement learning framework that integrates a Temporal Graph Neural Network (TimeGNN) to model multidimensional temporal states and introduces a discrete-continuous hybrid action space within a multi-agent deep deterministic policy gradient algorithm. This approach jointly optimizes fine-grained task partitioning, transmission power allocation, and scheduling priorities. Experimental results demonstrate that the proposed method significantly accelerates policy convergence, improves energy efficiency and latency performance, enhances task completion rates, and exhibits strong scalability and practical applicability.

0 citationsRead paper

Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey

Sep 29, 2025

A systematic survey of multimodal large language models (MLLMs) for multimodal sentiment recognition and reasoning remains absent. Method: This work presents the first comprehensive review of this interdisciplinary domain, covering mainstream MLLM architectures, multimodal sentiment datasets, benchmark performance, and technical pathways for cross-modal fusion and semantic reasoning. Through rigorous literature screening and horizontal comparative analysis, we identify core challenges—including insufficient model robustness, difficulties in cross-modal alignment, and weak fine-grained sentiment modeling. Contribution/Results: We propose promising future directions: enhanced interpretability, dynamic modality weighting, and causal sentiment reasoning. Concurrently, we release an open-source resource repository—including curated models, datasets, and evaluation code—to provide researchers with an authoritative survey, practical guidelines, and reproducible benchmarks—thereby advancing MLLM-driven affective intelligence.

0 citationsRead paper