A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究揭示了模型合并中的安全风险,提出Basin-Aware Jailbreak方法生成对抗性后缀,以解决共享预训练基础模型的合并模型家族面临的普遍越狱威胁。
📝 Abstract
Model merging enables combining multiple fine-tuned models without additional training, but its safety implications remain poorly understood. Prior work primarily attributes merging risks to unsafe constituent models, implicitly assuming that merging individually aligned models preserves safety. In contrast, we show that model merging reveals a previously overlooked jailbreak risk rooted in the pretrained foundation model, even when all constituent models are individually safety-aligned. Motivated by this observation, we study a new threat setting where an attacker constructs jailbreak prompts that generalize across merged models sharing the same pretrained backbone, without access to the exact merging coefficients or constituent checkpoints. To exploit this phenomenon, we propose \textbf{Basin-Aware Jailbreak (BAJ)}, which formulates jailbreak generation as a min--max optimization over the merging space to produce transferable adversarial suffixes across merged model families. Experiments across diverse backbones and merging settings show that BAJ achieves consistently high transfer success rates and remains effective under existing defenses.
Problem

Research questions and friction points this paper is trying to address.

model merging
jailbreak risk
pretrained foundation model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Basin-Aware Jailbreak
model merging
min-max optimization
adversarial suffixes
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yu Zhe
Yu Zhe
RIKEN AIP
Adversarial Machine Learning
Y
Yixin Tan
Institute of Science Tokyo
J
Junhao Wei
RIKEN AIP, Institute of Science Tokyo
C
Chen Wang
Zhejiang University