Manipulating Feature Visualizations with Gradient Slingshots
This work exposes a critical credibility vulnerability in feature visualization (FV) for deep neural network interpretability: FV outputs are susceptible to stealthy manipulation, leading to erroneous attribution of neuron semantics. To address this, we propose the first model-architecture-agnostic targeted FV manipulation method. Our approach integrates gradient redirection (via Slingshot optimization), adversarial latent-space perturbations, and neuron-activation-constrained regularization to achieve “semantic masking”—i.e., seamless substitution of a target neuron’s original FV explanation with an arbitrary user-specified semantic concept. Experiments across CNNs and Vision Transformers demonstrate successful concealment of functionally critical neurons: model accuracy degrades by less than 0.3%, yet FV-based auditing yields a 92% false-negative rate in detecting manipulated neurons. These results underscore the fragility of prevailing FV techniques and establish a new paradigm for robust model auditing and interpretability governance.