Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Editing Jailbreaks

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Open-weight large language models face a low-cost white-box threat from representation engineering attacks. Attackers can estimate refusal directions and search for projection-matrix edits that suppress safety alignment while preserving general capabilities, within minutes on a single GPU and without gradient-based training. We propose Bait-and-Recover, a weight-level defense that places a bait adapter where attackers read activations and a paired recovery adapter at the subsequent layer. Trained via gradient routing, this decouples the observation path from the behavior path. By actively poisoning the residual signal used for measurement, Bait-and-Recover disrupts the attacker's edit search, while the recovery layer restores clean downstream computation. Across four open-weight models, our defense raises the minimum refusal rate against white-box edit searches from 16.25% to 71.75% under a strict behavior-preservation budget (KL<= 0.10), with negligible impact on general benchmarks. By invalidating the core measurement assumption of these attacks, observation-path poisoning offers a practical complement to behavior-level safety training.
Problem

Research questions and friction points this paper is trying to address.

large language models
white-box threat
representation engineering attacks
refusal directions
projection-matrix edits
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bait-and-Recover
weight-level defense
gradient routing
observation-path poisoning
residual signal
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tian Gao
Anhui Laboratory for Safe Artificial Intelligence in the Yangtze River Delta
Zhipeng Xie
Zhipeng Xie
School of Computer Science, Fudan University
Natural Language ProcessingMachine LearningData MiningBioinformatics
Y
Yuhao Wu
Anhui Laboratory for Safe Artificial Intelligence in the Yangtze River Delta
Junhua Liu
Junhua Liu
University of Southern California
Multimedia SystemsVR/AR/XRAI/ML Systems
Xin Fang
Xin Fang
Emory University
MicrobiologyAntibioitics