SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models

📅 2026-05-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of defending multimodal large language models (MLLMs) against jailbreak attacks, which is hindered by input heterogeneity and the limitations of existing approaches that rely on costly fine-tuning or post-processing, often suffering from poor generalization and performance trade-offs. The authors propose SafeSteer, a novel mechanism that, for the first time, reveals the inherent safety discrimination capability of MLLMs during decoding. By introducing a lightweight decoding probe and cross-modal semantic alignment vectors, SafeSteer dynamically detects and corrects harmful outputs without requiring model fine-tuning. The method effectively transfers textual safety alignment to the visual modality, achieving up to a 33.40% improvement in safety performance across multiple mainstream MLLMs while preserving model utility, thereby establishing an efficient balance between safety and functionality.
📝 Abstract
Multimodal large language models (MLLMs) are gaining increasing attention. Due to the heterogeneity of their input features, they face significant challenges in terms of jailbreak defenses. Current defense methods rely on costly fine-tuning or inefficient post-hoc interventions, limiting their ability to address novel attacks and involving performance trade-offs. To address the above issues, we explore the inherent safety capabilities within MLLMs and quantify their intrinsic ability to discern harmfulness at decoding stage. We observe that 1) MLLMs can distinguish the harmful and harmless inputs during decoding process, 2) Image-based attacks are more stealthy. Based on these insights, we introduce SafeSteer, a decoding-level defense mechanism for MLLMs. Specifically, it includes a Decoding-Probe, a lightweight probe for detecting and correcting harmful output during decoding, which iteratively steers the decoding process toward safety. Furthermore, a modal semantic alignment vector is integrated to transfer the strong textual safety alignment to the vision modality. Experiments on multiple MLLMs demonstrate that SafeSterr can improve MLLMs' safety by up to 33.40\% without fine-tuning. Notably, it can maintain the effectiveness of MLLMs, ensuring a balance between their helpfulness and harmlessness.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Jailbreak Defense
Decoding-level Safety
Harmful Output Detection
Modal Heterogeneity
Innovation

Methods, ideas, or system contributions that make the work stand out.

decoding-level defense
multimodal large language models
safety alignment
Decoding-Probe
modal semantic alignment
💼 Related Jobs
No related jobs found.
Xinyi Zeng
Xinyi Zeng
Sichuan University
Medical Image SegmentationMedical Image ReconstructionMulti-modal Learning
X
Xue Yang
Shanghai Jiao Tong University
Jingyuan Zhang
Jingyuan Zhang
Kuaishou
Natural Language ProcessingLarge Language Model
H
Huanqian Yan
School of Computer Science and Technology, Beihang University
Xiang Chen
Xiang Chen
Nanjing University of Science and Technology
Computer VisionImage ProcessingArtificial IntelligenceDeep Learning
K
Kaiwen Wei
Chongqing University
H
Hankun Kang
Wuhan University
Y
Yu Tian
Tsinghua University, Beijing, China