Do VLMs Share Safety Neurons Across Modalities?

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过因果神经元级分析,揭示视觉-语言模型中安全机制的问题,并提出两阶段检测管道和两个基准来解决跨模态的安全性问题。
📝 Abstract
Vision-language models (VLMs) can comply with harmful requests delivered through images, even when their LLM backbones would refuse the same content in text. While prior work characterizes these jailbreaks empirically or at the representation level, how visual inputs perturb safety pathways at the neuron level remains uncharted. We close this gap with a causal, neuron-level analysis of safety mechanisms in 10 VLMs. We propose a two-stage detection pipeline with iterative ablation that accounts for self-repair, and introduce two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval, which decouple visual and textual safety signals. Our analysis reveals: (i) Text safety in VLMs is localizable: $\sim$88 neurons ($<$0.01%) whose targeted ablation substantially reduces refusal. (ii) Text safety neurons constitute the dominant refusal pathway: ablating them is the only intervention that consistently and substantially reduces refusal across all models. (iii) Visual safety is high-dimensional and diffuse at the single-neuron level: text safety concentrates in $\sim$5 subspace directions while visual safety requires $\geq$50. This gap holds across architectures, explaining why current alignment has not closed the visual safety gap. Project page is at: https://jiaxuan-li.github.io/vlm-safety-neuron/ Warning: this paper may include examples of harmful content.
Problem

Research questions and friction points this paper is trying to address.

Vision-language models
Safety neurons
Visual inputs
Neuron-level analysis
Safety mechanisms
Innovation

Methods, ideas, or system contributions that make the work stand out.

causal neuron-level analysis
two-stage detection pipeline
modality-isolated benchmarks
safety mechanisms
J
Jiaxuan Li
SB Intuitions Corp., Tokyo, Japan
J
Jiahao Zhang
SB Intuitions Corp., Tokyo, Japan
D
Duc Minh Vo
SB Intuitions Corp., Tokyo, Japan
Huy H. Nguyen
Huy H. Nguyen
SB Intuitions
Machine LearningResponsible AI
Pride Kavumba
Pride Kavumba
Tohoku University
natural language processingartificial intelligence
Koki Wataoka
Koki Wataoka
SB Intuitions
AI SafetyResponsible AILLM