RAFM-SER++: A Lightweight Multimodal Emotion Recognition Framework for Real-Time Behavioral Monitoring in Surveillance Systems

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决计算成本高限制实时监控系统部署的问题,提出RAFM-SER++框架,采用轻量级单向残差注意力机制融合多模态信息,实现高效情绪识别。
📝 Abstract
Recent multimodal Speech Emotion Recognition (SER) systems achieve high accuracy through interaction-heavy cross-modal transformers, but their computational cost limits deployment in latency-sensitive and resource-constrained surveillance systems. To address this challenge, we propose RAFM_SER++, a lightweight multimodal SER framework featuring an asymmetric Residual Attention Fusion Mechanism (RAFM). Rather than relying on computationally expensive bidirectional interactions, RAFM injects affective speech cues into semantic text representations through a one-directional residual attention pathway. Combined with a BYOL-inspired cross-modal alignment objective and attention-guided pooling, the proposed framework improves multimodal representation learning while maintaining low computational overhead. Experiments on the IEMOCAP and ESD benchmarks demonstrate that RAFM_SER++ consistently outperforms the HuBERT-Base baseline and achieves a superior accuracy-efficiency trade-off compared with the state-of-the-art MemoCMT. Specifically, RAFM_SER++ reduces trainable parameters by more than 60%, achieves faster inference (79.60 it/s), and attains BACC scores of 81.10% on IEMOCAP and 95.39% on ESD. These results indicate that lightweight asymmetric multimodal fusion is an effective alternative to interaction-heavy cross-modal transformers for real-time surveillance applications.
Problem

Research questions and friction points this paper is trying to address.

multimodal Speech Emotion Recognition
latency-sensitive
resource-constrained
computational cost
surveillance systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lightweight Multimodal Emotion Recognition
Asymmetric Residual Attention Fusion Mechanism
Cross-modal Alignment
Attention-guided Pooling
Real-time Surveillance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Ngo Truong Dinh
Department of Data Science, Industrial University of HCM City, Ho Chi Minh City, Vietnam
T
Tung-Lam Bui
Le Hong Phong High School for the Gifted, Ho Chi Minh City, Vietnam
C
Chi-Trung Duong
Department of Data Science, Industrial University of HCM City, Ho Chi Minh City, Vietnam
V
Vien Nguyen Thi
Department of Data Science, Industrial University of HCM City, Ho Chi Minh City, Vietnam
V
Viet-Anh Nguyen
Faculty of Information Technology, FPT University, Ho Chi Minh City, Vietnam
Phuc-Lu Le
Phuc-Lu Le
University of Science, VNU-HCM
Graph theoryPrivacy PreservingInformation theory