CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection

πŸ“… 2026-07-13
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the prevalent issue of large language models over-predicting sarcasm in zero-shot settings and the limited cross-lingual transferability of prosodic cues. To mitigate these challenges, the authors propose a training-free multimodal framework that integrates a novel bidirectional symmetric prompt calibration mechanism (BiCAL) with an acoustic late-fusion rescue strategy (ALFR), dynamically weighting prosodic features to correct model bias. The approach achieves a zero-shot text-only Macro-F1 of 0.787 on MUStARD and yields improvements of up to +0.382 over weak baselines on CMMA (p < 10⁻⁴³), significantly outperforming existing methods while offering strong interpretability and cross-lingual generalization capabilities.
πŸ“ Abstract
Sarcasm detection, the identification of discrepancies between literal and intended meaning, is a fundamental task in affective computing. However, zero-shot instruction-tuned Large Language Models (LLMs) systematically over-predict the positive (sarcastic) class across the entire capability spectrum, while the prosodic cues humans rely on remain underexploited and transfer unevenly across languages. We introduce CHARM (Charge Calibration and Acoustic Rescue for Multimodal Sarcasm Detection), a training-free framework that couples two modules. Bidirectional Charge Calibration (BiCAL) steers the LLM toward opposing sarcastic and literal verdicts along a symmetric axis of charged prompts; the induced directional biases cancel by construction, and a simple aggregation recovers an unbiased pragmatic signal. Acoustic Late-Fusion Rescue (ALFR) then fuses the calibrated votes with prosodic descriptors and LLM-generated auditory-perception probes through a shallow classifier, actively down-weighting saturated text votes in favour of acoustic evidence. Without fine-tuning any backbone, BiCAL attains the highest reported zero-shot text-only Macro-F1 of 0.787 on MUStARD, while ALFR lifts weak backbones by up to +0.382 Macro-F1 on CMMA. A Stouffer meta-analysis confirms statistical significance on MUStARD and CMMA (Z = 13.89 and Z = 34.64, respectively; p < 10^-43). Our analysis further uncovers a cross-cultural prosodic decoupling: low-level acoustics fail to transfer across languages, whereas high-level perceptual abstractions remain robust. Together, these components yield an explainable, cross-lingual multimodal detector.
Problem

Research questions and friction points this paper is trying to address.

sarcasm detection
zero-shot LLMs
prosodic cues
cross-lingual transfer
class imbalance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bidirectional Charge Calibration
Acoustic Late-Fusion Rescue
zero-shot sarcasm detection
multimodal fusion
cross-lingual prosody
Qiyang Sun
Qiyang Sun
Imperial College London
Yi Chang
Yi Chang
Imperial College London
Affective ComputingComputer AuditionDigital Health
Y
Yupei Li
GLAM – the Group on Language, Audio, & Music, Imperial College London, UK
Xi Shao
Xi Shao
Professor of Computer Engineering,Nanjing University of Posts and Telecommunications
Multimedia Information SystemComputer Audition
Zixing Zhang
Zixing Zhang
Professor, Hunan University
Artifical IntelligenceSpeech ProcessingAffective ComputingDigital HealthAutomatic Speech Recognition
B
BjΓΆrn W. Schuller
GLAM – the Group on Language, Audio, & Music, Imperial College London, UK; CHI – Chair of Health Informatics, TUM University Hospital, Germany; relAI – the Konrad Zuse School of Excellence in Reliable AI, Germany; MDSI – Munich Data Science Institute, Germany; MCML – Munich Center for Machine Learning, Germany