Interpreting Object-Dependent Concept Brittleness in Text-to-Image Diffusion Models

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对文本到图像扩散模型中对象依赖的概念脆弱性问题,提出了一种基于稀疏自编码器空间的可解释性框架,并引入轻量级推理时校正策略来解决。
📝 Abstract
Although text-to-image diffusion models generally exhibit strong prompt-following ability, we identify a persistent and previously underexplored failure pattern in which a small subset of prompts differing only in the object consistently fails to realize the same target concept under identical generation settings. We term this phenomenon object-dependent concept brittleness. Such cases suggest systematic internal blind spots rather than random sampling noise. In this paper, we present an interpretability-oriented framework to audit and minimally correct these failures. Our key idea is to analyze denoising trajectories in a step-wise sparse autoencoder (SAE) space, where abstract style and attribute concepts become more separable than in the raw denoising representation. This sparse space enables us to compare successful and failed generations, identify concept dimensions whose evidence is missing, weakened, or temporally delayed, and construct class-level concept prototypes from reliable class-consistent samples. Based on this audit process, we introduce a lightweight inference-time correction strategy that interpolates denoising features toward the corresponding prototype in SAE space. Rather than serving as a task-specific retraining method, this intervention acts as a validation of the diagnosed concept deficiency. We evaluate the proposed framework on style and attribute failure cases across multiple diffusion backbones, with significant improvements in concept consistency, text fidelity, and repair success. Further analyses show that deeper denoising representations provide clearer concept structure, while early-stage intervention offers the strongest correction leverage. Code is available at https://github.com/Metecade/Object-Dependent-Concept-Brittleness.
Problem

Research questions and friction points this paper is trying to address.

object-dependent concept brittleness
text-to-image diffusion models
prompt-following ability
Innovation

Methods, ideas, or system contributions that make the work stand out.

object-dependent concept brittleness
sparse autoencoder (SAE) space
denoising trajectories
inference-time correction
concept consistency
🔎 Similar Papers
Y
Yifan Yuan
School of Artificial Intelligence, Shenzhen University, Shenzhen, China
X
Xiangyu Liu
School of Mathematical Sciences, Shenzhen University, Shenzhen, China
Hongming Shan
Hongming Shan
Fudan University; Rensselaer Polytechnic institute
Machine LearningMedical ImagingComputer Vision
Yu Han
Yu Han
Professor of Chemistry, South China University of Technology
NanomaterialsElectron MicroscopyCatalysis
Y
Yu Jiang
Department of Statistics and Data Science, National University of Singapore, Singapore, Singapore
Hao Tan
Hao Tan
Adobe Research
Vision and Language3D Multimodal
J
Junping Zhang
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China
Linlin Shen
Linlin Shen
Shenzhen University
Deep LearningComputer VisionFacial Analysis/RecognitionMedical Image Analysis