Institution profile

Rider University

Academic institutionnorthamerica · us
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework

May 24, 2025

Speech-driven multimodal large language models (e.g., SpeechGPT) exhibit alignment vulnerabilities unique to the speech modality—stemming from temporal characteristics of speech, phonetic variability, and ASR uncertainty—enabling adversaries to bypass safety guardrails. Method: We propose the first white-box adversarial attack framework targeting speech tokenizers: by reverse-engineering the speech tokenizer, we perform token-level perturbation optimization directly in the speech embedding space to synthesize playable adversarial audio—eliminating reliance on black-box TTS or manual construction. Contribution/Results: Our method achieves an 89% attack success rate on SpeechGPT, substantially outperforming existing speech jailbreaking approaches. It is the first to expose structural weaknesses in speech-modal alignment mechanisms, offering a novel paradigm for security evaluation and robust training of multimodal foundation models.

0 citationsRead paper
Recent publications

Latest Papers

Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework

May 24, 2025

Speech-driven multimodal large language models (e.g., SpeechGPT) exhibit alignment vulnerabilities unique to the speech modality—stemming from temporal characteristics of speech, phonetic variability, and ASR uncertainty—enabling adversaries to bypass safety guardrails. Method: We propose the first white-box adversarial attack framework targeting speech tokenizers: by reverse-engineering the speech tokenizer, we perform token-level perturbation optimization directly in the speech embedding space to synthesize playable adversarial audio—eliminating reliance on black-box TTS or manual construction. Contribution/Results: Our method achieves an 89% attack success rate on SpeechGPT, substantially outperforming existing speech jailbreaking approaches. It is the first to expose structural weaknesses in speech-modal alignment mechanisms, offering a novel paradigm for security evaluation and robust training of multimodal foundation models.

0 citationsRead paper