Safe-Control: A Safety Patch for Mitigating Unsafe Content in Text-to-Image Generation Models
Existing text-to-image (T2I) models are vulnerable to misuse for generating unsafe content, while mainstream safety mechanisms exhibit poor robustness under distribution shifts or adversarial attacks and typically require model fine-tuning. To address this, we propose a plug-and-play safety patching framework that operates without modifying the original model’s weights. Our method introduces learnable, data-driven safety-aware conditioning signals into intermediate layers of the frozen diffusion model during the denoising process. It supports multi-strategy fusion and cross-model transferability, significantly enhancing resilience against both distribution shifts and adversarial prompts. Evaluated on six state-of-the-art T2I models, our approach reduces unsafe image generation to 7%, outperforming seven existing SOTA safety methods, while preserving high-fidelity image quality and strong text–image alignment.