Talking Like a Phisher: LLM-Based Attacks on Voice Phishing Classifiers

📅 2025-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Vishing detection models are vulnerable to semantic-preserving adversarial attacks, yet existing studies lack systematic evaluation of large language model (LLM)-driven, highly stealthy, and low-cost attacks. To address this gap, we propose the first LLM-based adversarial text generation framework leveraging prompt engineering and semantic masking. Using four commercial LLMs—including GPT-4o—we generate semantically consistent, fraud-intent-preserving adversarial transcriptions on the Korean vishing dataset KorCCViD. Experiments show that GPT-4o-generated samples reduce accuracy of state-of-the-art classifiers by 30.96%, with average generation time under 9 seconds and negligible cost. BERTScore confirms high semantic fidelity (≥0.87). This work provides the first empirical evidence that off-the-shelf LLMs can efficiently execute semantic-preserving adversarial attacks against vishing detectors, exposing a fundamental robustness deficiency in current systems and establishing a critical benchmark for developing resilient defenses.

Technology Category

Application Category

📝 Abstract
Voice phishing (vishing) remains a persistent threat in cybersecurity, exploiting human trust through persuasive speech. While machine learning (ML)-based classifiers have shown promise in detecting malicious call transcripts, they remain vulnerable to adversarial manipulations that preserve semantic content. In this study, we explore a novel attack vector where large language models (LLMs) are leveraged to generate adversarial vishing transcripts that evade detection while maintaining deceptive intent. We construct a systematic attack pipeline that employs prompt engineering and semantic obfuscation to transform real-world vishing scripts using four commercial LLMs. The generated transcripts are evaluated against multiple ML classifiers trained on a real-world Korean vishing dataset (KorCCViD) with statistical testing. Our experiments reveal that LLM-generated transcripts are both practically and statistically effective against ML-based classifiers. In particular, transcripts crafted by GPT-4o significantly reduce classifier accuracy (by up to 30.96%) while maintaining high semantic similarity, as measured by BERTScore. Moreover, these attacks are both time-efficient and cost-effective, with average generation times under 9 seconds and negligible financial cost per query. The results underscore the pressing need for more resilient vishing detection frameworks and highlight the imperative for LLM providers to enforce stronger safeguards against prompt misuse in adversarial social engineering contexts.
Problem

Research questions and friction points this paper is trying to address.

Evading ML-based voice phishing detection using LLM-generated transcripts
Assessing adversarial vishing attacks on Korean phishing classifiers
Evaluating cost-effective LLM manipulations for bypassing security systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-generated adversarial vishing transcripts evade detection
Prompt engineering and semantic obfuscation transform vishing scripts
GPT-4o reduces classifier accuracy by 30.96%
Universiti Sains Malaysia
W
Wenhao Li
Cybersecurity Research Centre, Universiti Sains Malaysia, Pulau Pinang, Malaysia
S
Selvakumar Manickam
Cybersecurity Research Centre, Universiti Sains Malaysia, Pulau Pinang, Malaysia
Y
Yung-wey Chong
School of Computer Sciences, Universiti Sains Malaysia, Pulau Pinang, Malaysia
Shankar Karuppayah
Shankar Karuppayah
Deputy Director and Senior Lecturer at Universiti Sains Malaysia
BotnetsCyber SecurityPeer-to-peer networks