Certifying Concept Unlearning in Text-to-Image Diffusion Models

๐Ÿ“… 2026-09-10
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บไธ€็งๆ–ฐ็š„่ฎค่ฏๆก†ๆžถ๏ผŒ้€š่ฟ‡็ปŸ่ฎก่ฎค่ฏๅ’Œๆœ€ๅๆƒ…ๅ†ตๅˆ†ๆžๆฅ้‡ๅŒ–ๆ–‡ๆœฌๅˆฐๅ›พๅƒๆ‰ฉๆ•ฃๆจกๅž‹ไธญๆฆ‚ๅฟต้—ๅฟ˜็š„ๆฎ‹ไฝ™ๆณ„ๆผ้ฃŽ้™ฉใ€‚
๐Ÿ“ Abstract
Existing evaluations of concept unlearning in text-to-image (T2I) diffusion models primarily rely on attack success rates obtained through automated adversarial prompt search. However, these metrics provide only empirical evidence over a finite set of queries and leave residual leakage over the broader prompt space largely unquantified. This limitation can lead to overestimating unlearning effectiveness and underestimating safety risks. To address this gap, we introduce a novel certification framework for T2I concept unlearning that provides high-confidence guarantees with bounded error on residual concept leakage. Our approach combines statistical certification with worst-case analysis along concept-relevant embedding directions to derive explicit upper bounds on leakage probability under user-specified confidence levels. We evaluate our framework across three major concept categories namely NSFW content, artistic styles, and celebrity identities, and six state-of-the-art unlearning methods. Certified leakage bounds consistently exceed standard attack success rates by 16.2%, uncovering substantial residual risks missed by existing evaluation protocols. Crucially, our results demonstrate that empirical attack-based evaluations can significantly underestimate residual leakage and establish certification as a necessary complement for reliable auditing of concept unlearning in T2I diffusion models.
Problem

Research questions and friction points this paper is trying to address.

concept unlearning
text-to-image diffusion models
residual leakage
adversarial prompt search
safety risks
Innovation

Methods, ideas, or system contributions that make the work stand out.

certification framework
concept unlearning
residual leakage
statistical certification
worst-case analysis