Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models

๐Ÿ“… 2026-08-17
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
็ ”็ฉถๆๅ‡บไบ†ไธ€็งๅไธบ'ๆ„šไบบ้‡‘'็š„้˜ฒๅพกๆ–นๆณ•๏ผŒ้€š่ฟ‡ๅœจๆจกๅž‹ๅ—ๅˆฐๆ”ปๅ‡ปๆ—ถ็”Ÿๆˆ่™šๅ‡ๅ“ๅบ”ๆฅๆฌบ้ช—ๆ”ปๅ‡ป่€…๏ผŒไปŽ่€ŒไฟๆŠคๅผ€ๆ”พๆƒ้‡ๆจกๅž‹็š„ๅฎ‰ๅ…จๆ€งใ€‚
๐Ÿ“ Abstract
Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defense we are aware of prevents it durably. What cannot be prevented can be deceived. Our defense, decoy hardening ("Fool's Gold"), concedes the refusal strip and poisons its payoff: once refusal is stripped, most answers to hazardous operational requests are confident, fluent decoys whose critical elements are falsified. Decoys are trained inside a differentiable simulation of the attack, expressing only in the attacked state; a refusal pin and benign leash hold clean-state behavior to the original. We instantiate it on seven models from five families (9B-122B, dense and mixture-of-experts). On the six models passing our pre-registered efficacy gate, 0.51-0.90 of attacked-state responses to held-out prompts are decoys, +0.27-0.84 attributable to the defense; all six stay within registered benign-behavior and capability budgets; the seventh (smaller) fails the gate (boundary case). Rates replicate on a frozen test split or untouched strata. The claim is epistemic: without independent ground truth, no observation surface we tested separates falsified answers from correct ones - on external red-team benchmarks' CBRNE-adjacent slice, the defended 122B is fatally wrong on 0.82-0.86 of matched-quality answers vs at most 0.10 undefended. Repeated sampling does not restore trust: element-wise consensus at K=64 reconstructs a fully usable procedure on 0.083-0.625 of prompts where the instrument validates, vs 0.58-0.96 undefended, with no label-free way to tell the regimes apart; on the weakest such model the claim is per-draw only. We evaluate chemical and biological hazards; the defense does not address in-context jailbreaks and protects only the initially released defended weights.
Problem

Research questions and friction points this paper is trying to address.

Safety-Removal Attacks
Open-Weight Models
Defensive Deception
Innovation

Methods, ideas, or system contributions that make the work stand out.

Defensive Deception
Safety-Removal Attacks
Decoy Hardening
Open-Weight Models
๐Ÿ”Ž Similar Papers
2024-08-01arXiv.orgCitations: 20
2024-07-01Conference on Empirical Methods in Natural Language ProcessingCitations: 2