GuardPaint:SpeculativeSafetyDecodingforText-to-ImageGeneration

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决文本到图像生成中的安全问题,GuardPaint通过在扩散过程中监控和修复不安全区域,而无需修改基础模型,减少有害内容生成。
📝 Abstract
Text-to-image (T2I) diffusion models offer powerful visual generation, but their controllability creates a critical safety challenge: adversarial prompts can steer the denoising trajectory toward policy-violating content such as explicit nudity or graphic violence. Existing safeguards mostly act before generation through prompt filtering or after generation through image classification, leaving the diffusion process itself unguarded and often yielding only refusal rather than safe visual repair. We introduce GuardPaint, a speculative decoding framework for safe T2I generation that intervenes inside the diffusion trajectory without modifying the base model. A lightweight auditor monitors intermediate images, localizes unsafe regions, and triggers surgical inpainting repair only where needed. Candidate repairs are generated by a policy-aligned inpainter and selected through a guarded tournament that accepts edits only when they improve policy compliance while preserving prompt fidelity and perceptual quality. Across five jailbreak families SneakPrompt, MMA, PGJ, DACA, and RABell and UNet/flow-matching models including SD~1.5, SDXL, SD~3.5, and FLUX.1-dev. GuardPaint reduces attack success and harmful generations with minimal degradation to image quality, prompt fidelity, and benign behavior. Content warning: This paper contains examples involving nudity and violence that some readers may find disturbing, distressing, or offensive.
Problem

Research questions and friction points this paper is trying to address.

Text-to-Image
Diffusion Models
Safety Challenge
Adversarial Prompts
Policy-Violating Content
Innovation

Methods, ideas, or system contributions that make the work stand out.

speculative decoding
diffusion trajectory intervention
surgical inpainting repair
guarded tournament
💼 Related Jobs
No related jobs found.
S
Shreyash Dhoot
Pragya Lab, BITS Pilani Goa, India
P
Paras Dhiman
Pragya Lab, BITS Pilani Goa, India
A
Arsh Abbas Naqvi
Pragya Lab, BITS Pilani Goa, India
A
Aranbi Dutta
Pragya Lab, BITS Pilani Goa, India
Aman Chadha
Aman Chadha
GenAI Leadership @ Apple • Stanford AI • UW-Madison ECE • Ex: Apple, AWS, Alexa, Nvidia
Multimodal AINatural Language ProcessingComputer VisionSpeech ProcessingRecommender Systems
Vinija Jain
Vinija Jain
Meta | Ex: Amazon, Oracle, Palo Alto Networks
AINatural Language ProcessingMultimodal AIRecommender SystemsInformation Retrieval
A
Amitava Das
Pragya Lab, BITS Pilani Goa, India