A Dual-Purpose Framework for Backdoor Defense and Backdoor Amplification in Diffusion Models

📅 2025-02-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Diffusion models are vulnerable to backdoor attacks—maliciously triggering harmful content generation upon injection of specific perturbations—yet existing defense and attack analysis methods remain fragmented and suboptimal. This paper proposes PureDiffusion, the first unified framework enabling *simultaneous* high-robustness backdoor detection and controllable attack enhancement. Its core innovation is a dual-loss trigger inversion reconstruction mechanism grounded in temporal distribution shift modeling and denoising consistency constraints, supporting bidirectional optimization. On the defense side, it achieves ≈100% detection accuracy—substantially surpassing state-of-the-art methods. On the attack side, lightweight trigger reinforcement training boosts attack success rates to ≈100% while reducing training time by 20×. PureDiffusion establishes a novel, interpretable, and reusable paradigm for backdoor security research in diffusion models.

Technology Category

Application Category

📝 Abstract
Diffusion models have emerged as state-of-the-art generative frameworks, excelling in producing high-quality multi-modal samples. However, recent studies have revealed their vulnerability to backdoor attacks, where backdoored models generate specific, undesirable outputs called backdoor target (e.g., harmful images) when a pre-defined trigger is embedded to their inputs. In this paper, we propose PureDiffusion, a dual-purpose framework that simultaneously serves two contrasting roles: backdoor defense and backdoor attack amplification. For defense, we introduce two novel loss functions to invert backdoor triggers embedded in diffusion models. The first leverages trigger-induced distribution shifts across multiple timesteps of the diffusion process, while the second exploits the denoising consistency effect when a backdoor is activated. Once an accurate trigger inversion is achieved, we develop a backdoor detection method that analyzes both the inverted trigger and the generated backdoor targets to identify backdoor attacks. In terms of attack amplification with the role of an attacker, we describe how our trigger inversion algorithm can be used to reinforce the original trigger embedded in the backdoored diffusion model. This significantly boosts attack performance while reducing the required backdoor training time. Experimental results demonstrate that PureDiffusion achieves near-perfect detection accuracy, outperforming existing defenses by a large margin, particularly against complex trigger patterns. Additionally, in an attack scenario, our attack amplification approach elevates the attack success rate (ASR) of existing backdoor attacks to nearly 100% while reducing training time by up to 20x.
Problem

Research questions and friction points this paper is trying to address.

Defends against backdoor attacks in diffusion models
Amplifies backdoor attack effectiveness in diffusion models
Inverts and detects backdoor triggers in diffusion models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-purpose framework for defense and amplification
Novel loss functions for trigger inversion
Backdoor detection via inverted trigger analysis
🔎 Similar Papers
No similar papers found.
V
Vu Tuan Truong Long
INRS, University of Qu´ebec, Montr´eal, QC H5A 1K6, Canada
B
Bao Le
INRS, University of Qu´ebec, Montr´eal, QC H5A 1K6, Canada