🤖 AI Summary
This study investigates whether large audio language models encode cross-lingual emotion through language-agnostic neural mechanisms. To this end, the authors propose a framework for identifying multilingual emotion neurons (MLENs), integrating neuron-level interpretability analysis, causal intervention, and a consistency regularization fusion (CR-Fusion) strategy to jointly discover emotion-sensitive neurons across twelve languages. The work provides the first neuron-level causal evidence of cross-lingual emotion encoding in large audio language models and reveals the non-redundant contributions and asymmetric gains of low-resource languages in affective transfer. Experiments demonstrate that MLENs exhibit high precision and strong transferability across four prominent models, significantly outperforming monolingual approaches in zero-shot and low-resource settings.
📝 Abstract
Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks, yet it remains unclear whether they encode emotion through language-specific correlations or language-agnostic representations. We present the first neuron-level interpretability study of this question. We define Multilingual Emotion Neurons (MLENs) as functional units exhibiting stable emotional selectivity and aligned causal effects across languages, and introduce Consistency-Regularized Fusion (CR-Fusion) to identify them. Across four modern LALMs and 12 typologically diverse languages, emotion-sensitive neurons identified independently per language show minimal overlap, and additional monolingual identification data saturates quickly without isolating more transferable units, motivating identification from pooled cross-lingual evidence. Causal interventions demonstrate that MLENs identified by CR-Fusion provide more precise and transferable affective control than monolingual neuron sets in both zero-shot and low-resource settings. Leave-one-out ablations further reveal asymmetric transfer: individual identification languages, including low-resource ones, contribute non-redundant evidence, while several low-resource languages benefit most from the resulting cross-lingual transfer. Together, our findings provide the first causal, neuron-level account of how LALMs encode emotion across languages, and establish multilingual neuron identification as an effective mechanism for understanding cross-lingual affective behavior.