DAVE: Distribution-aware Attribution via ViT Gradient Decomposition
This work addresses the limitations of attribution maps generated by Vision Transformers (ViTs), which are often compromised by structured artifacts stemming from patch embeddings and attention mechanisms, yielding only coarse-grained and unstable block-level explanations. To overcome this, the authors propose a gradient decomposition method tailored to the architectural characteristics of ViTs, introducing— for the first time—distribution-aware modeling to mathematically disentangle the local, equivariant, and stable input–output mapping components from structural noise. This approach effectively suppresses architecture-induced artifacts, substantially enhancing both the stability and spatial resolution of attributions, and thereby producing high-fidelity, pixel-level explanation maps.