The Fragility of Jailbreak Robustness Across Operational States

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了单状态评估下越狱鲁棒性的问题,通过分析七个对齐模型和三种攻击在不同操作状态下的表现,揭示了系统提示改变对攻击成功率的影响。
📝 Abstract
Existing jailbreak evaluations typically characterize robustness using a single attack success rate (ASR) measured in a default configuration (the vanilla state). However, user-LLM interactions can induce diverse operational states beyond the vanilla state. In this work, we find that jailbreak robustness is highly fragile to operational-state variation: even when the attack remains fixed, changing only an ordinary system prompt not designed to affect safety can dramatically alter attack success rates. We systematically investigate this phenomenon across seven aligned models and three representative jailbreak attacks, observing substantial differences in ASR between vanilla and non-vanilla operational states. In one case, ASR increases by up to 56 percentage points (2% to 58%) solely due to a change in operational state. Remarkably, these increases occur even for attacks originally designed and optimized under vanilla-state evaluation. We further show that state-dependent robustness variation is systematically associated with differences in hidden representations along a refusal-related axis, and that projections onto this axis strongly predict jailbreak outcomes. Our results show that a single vanilla-state evaluation may not fully characterize jailbreak robustness, motivating evaluations that also examine how robustness changes across non-vanilla operational states.
Problem

Research questions and friction points this paper is trying to address.

jailbreak robustness
operational states
attack success rate
system prompt
Innovation

Methods, ideas, or system contributions that make the work stand out.

Operational-State Variation
Jailbreak Robustness
Hidden Representations
🔎 Similar Papers
No similar papers found.