Detecting Mental Manipulation in Speech via Synthetic Multi-Speaker Dialogue
This study addresses a critical gap in psychological manipulation detection by extending the task beyond text to the underexplored speech modality. The authors introduce SPEECHMENTALMANIP, the first benchmark for detecting covert manipulative language in speech, constructed by augmenting existing textual datasets with high-fidelity, speaker-consistent multi-speaker synthetic audio. Systematic evaluation combines few-shot audio-language models with human annotation to assess performance. Experimental results reveal that while models achieve high specificity in identifying manipulative utterances in speech, their recall is substantially lower than in the textual domain. Human evaluators also exhibit greater uncertainty, highlighting the inherent ambiguity of manipulative cues in spoken language and fundamental differences in cross-modal perception. These findings underscore the unique challenges posed by speech-based manipulation detection and call for modality-aware approaches.