Critic-Guided Reinforcement Unlearning in Text-to-Image Diffusion

📅 2026-01-06
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the problem of concept erasure in text-to-image diffusion models by proposing an efficient reinforcement learning–based removal method. By formulating the denoising process as a sequential decision-making problem, the approach introduces a timestep-aware critic and a CLIP-based noise-conditioned reward signal to enable policy gradient updates of the reverse diffusion kernel. This design significantly improves credit assignment accuracy and training stability. Empirical results demonstrate that the method achieves forgetting performance on multiple concepts that is either superior or comparable to strong existing baselines, while effectively preserving the model’s overall generation quality and prompt fidelity.

Technology Category

Application Category

📝 Abstract
Machine unlearning in text-to-image diffusion models aims to remove targeted concepts while preserving overall utility. Prior diffusion unlearning methods typically rely on supervised weight edits or global penalties; reinforcement-learning (RL) approaches, while flexible, often optimize sparse end-of-trajectory rewards, yielding high-variance updates and weak credit assignment. We present a general RL framework for diffusion unlearning that treats denoising as a sequential decision process and introduces a timestep-aware critic with noisy-step rewards. Concretely, we train a CLIP-based reward predictor on noisy latents and use its per-step signal to compute advantage estimates for policy-gradient updates of the reverse diffusion kernel. Our algorithm is simple to implement, supports off-policy reuse, and plugs into standard text-to-image backbones. Across multiple concepts, the method achieves better or comparable forgetting to strong baselines while maintaining image quality and benign prompt fidelity; ablations show that (i) per-step critics and (ii) noisy-conditioned rewards are key to stability and effectiveness. We release code and evaluation scripts to facilitate reproducibility and future research on RL-based diffusion unlearning.
Problem

Research questions and friction points this paper is trying to address.

machine unlearning
text-to-image diffusion
concept removal
reinforcement learning
diffusion models
Innovation

Methods, ideas, or system contributions that make the work stand out.

reinforcement unlearning
diffusion models
timestep-aware critic
noisy-step rewards
policy gradient
M
Mykola Vysotskyi
SoftServe
Z
Zahar Kohut
SoftServe
M
Mariia Shpir
SoftServe
T
Taras Rumezhak
SoftServe
V
Volodymyr Karpiv
SoftServe