Semantic Steering for Controllable Generation: Tuning-Free Concept Erasure in Multimodal Diffusion Transformers

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of sensitive content generation in multimodal diffusion Transformers (MM-DiTs), where existing concept erasure methods rely on model fine-tuning and are impractical for large-scale pretrained models. The study reveals, for the first time, that semantic representations in intermediate layers of MM-DiTs are most salient for concept retention. Building on this insight, the authors propose a training-free erasure strategy that analyzes block-level representations and injects a semantic steering vector into sparse tokens of the text branch, achieving precise control by operating only on early-to-intermediate layers. Evaluated across multiple MM-DiT architectures, the method demonstrates state-of-the-art concept erasure performance while offering high efficiency, broad applicability, adversarial robustness, and negligible computational overhead.
📝 Abstract
Multimodal Diffusion Transformers (MM-DiTs) have demonstrated remarkable text-to-image generation performance, surpassing traditional U-Net-based diffusion models. Nevertheless, their powerful generative capabilities also raise significant safety concerns, as they may generate sensitive or inappropriate content. While existing concept erasure methods aim to mitigate such risks, most require modifying model parameters, which are often architecture-specific and impractical for deployed larger models. Several tuning-free approaches face challenges when applied to advanced large-scale MM-DiTs due to their deeply embedded knowledge, broad semantic space, and context-dependent text encoders. To address these challenges, we propose to erase concepts by directly manipulating the model's internal representations. Our key insight, derived from an in-depth analysis of MM-DiT's block-wise generative roles, is that text-conditioned semantic representations are most salient in the middle blocks of MM-DiTs. Based on this, we extract representations of an unwanted concept and a desirable safe one from the middle block, construct a steering vector from their difference, and inject this single vector into consecutive early and middle blocks. By operating exclusively on the sparse text-branch tokens and leveraging the straight sampling trajectory of rectified flow, our method achieves effective concept erasure with negligible overhead and without any training. Extensive experiments across MM-DiT models demonstrate that our method achieves state-of-the-art performance in erasing diverse concepts, enables effective control over the final output, and remains robust to adversarial attacks.
Problem

Research questions and friction points this paper is trying to address.

concept erasure
multimodal diffusion transformers
controllable generation
safety
tuning-free
Innovation

Methods, ideas, or system contributions that make the work stand out.

concept erasure
semantic steering
multimodal diffusion transformers
tuning-free control
representation manipulation
💼 Related Jobs
No related jobs found.
Qiao Li
Qiao Li
Postdoctoral Research Fellow in IBME, Dept. Engineering Science, University of Oxford
Multi-dimensional Biomedical Signal ProcessingAdvanced Patient MonitoringArtifact and Noise AnalysisMachine LearningPhysiological Database
X
Xiaomeng Fu
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences
Y
Yuanshu Zhao
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences
Q
Qipeng Wang
Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences
J
Jiao Dai
Institute of Information Engineering, Chinese Academy of Sciences
J
Jizhong Han
Institute of Information Engineering, Chinese Academy of Sciences