Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种针对多模态大语言模型的黑盒越狱攻击方法TA-SPA,通过优化文本锚定的语义扰动来解决安全对齐问题。
📝 Abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language interaction, yet their safety alignment remains vulnerable to jailbreak attacks. A key challenge is that safety behavior learned in the textual space does not reliably transfer to fused cross-modal representations, leaving multimodal inputs exploitable through latent semantic cues. We propose Text-Anchored Semantic Perturbation Attack (TA-SPA), a black-box jailbreak framework that optimizes transferable perturbations in a text-anchored semantic space. TA-SPA integrates Text-Anchored Semantic Factorization (TASF), which encourages the separation of cross-modal semantic factors from modality-specific residuals, with Semantic-Preserving Augmentation (SPA), which diversifies harmful target anchors while preserving semantic consistency. Experiments show strong attack effectiveness and transfer to commercial MLLMs, with competitive performance under representative defenses. Additional controls and probing support the intended factorization without implying perfect disentanglement, motivating representation-level safety alignment beyond input-level filtering.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
jailbreak attacks
safety alignment
cross-modal representations
semantic cues
Innovation

Methods, ideas, or system contributions that make the work stand out.

Text-Anchored Semantic Perturbation Attack
TASF
Semantic-Preserving Augmentation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
W
Wenyun Li
Harbin Institute of Technology, Shenzhen, China; Pengcheng Laboratory, Shenzhen, China
Guiping Cao
Guiping Cao
PCL; SUSTech; CVTE Research; XJTU
Deep LearningComputer VisionMedical Image Processing
Xiangyuan Lan
Xiangyuan Lan
Pengcheng Laboratory
Multimodal LLMPlace RecognitionVisual TrackingPerson Re-identificationObject Detection
Z
Zheng Zhang
Harbin Institute of Technology, Shenzhen, China; Pengcheng Laboratory, Shenzhen, China