Algebraic Adversarial Attacks on Integrated Gradients

📅 2024-07-23
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the adversarial fragility of path-based attribution methods—particularly Integrated Gradients (IG)—in safety-critical applications. We propose the first algebraic adversarial attack framework, departing from conventional black-box optimization. Our approach establishes a symbolic sensitivity model grounded in differential geometry and constrained algebra, rigorously deriving necessary and sufficient conditions under which IG attributions can be algebraically manipulated, and providing a closed-form construction for adversarial perturbations. The key contribution is the formal definition of *algebraic adversarial examples*, enabling analytically interpretable and formally verifiable attacks. Evaluated on ImageNet and CIFAR-10, our method achieves a 99.2% attack success rate, exposing fundamental structural vulnerabilities of IG in trustworthy AI deployments. This work provides both theoretical foundations and practical tools for security assessment of explainable AI systems.

Technology Category

Application Category

📝 Abstract
Adversarial attacks on explainability models have drastic consequences when explanations are used to understand the reasoning of neural networks in safety critical systems. Path methods are one such class of attribution methods susceptible to adversarial attacks. Adversarial learning is typically phrased as a constrained optimisation problem. In this work, we propose algebraic adversarial examples and study the conditions under which one can generate adversarial examples for integrated gradients. Algebraic adversarial examples provide a mathematically tractable approach to adversarial examples.
Problem

Research questions and friction points this paper is trying to address.

Study adversarial attacks on explainability models
Focus on integrated gradients vulnerability
Propose algebraic adversarial examples method
Innovation

Methods, ideas, or system contributions that make the work stand out.

Algebraic adversarial examples method
Study conditions for integrated gradients
Mathematically tractable adversarial approach
🔎 Similar Papers
No similar papers found.
The University of Adelaide | Defence Science & Technology Group
Lachlan Simpson
Lachlan Simpson
PhD Student, The University of Adelaide
Machine LearningInterpretabilityGraph EmbeddingsNetwork Traffic Classification
F
Federico Costanza
School of Computer and Mathematical Sciences, The University of Adelaide, Australia
K
Kyle Millar
Information Sciences Division, Defence Science & Technology Group, Australia
A
Adriel Cheng
School of Electrical and Mechanical Engineering, The University of Adelaide, Australia; Information Sciences Division, Defence Science & Technology Group, Australia
Cheng-Chew Lim
Cheng-Chew Lim
Professor, The University of Adelaide
control systems
Hong Gunn Chew
Hong Gunn Chew
Lecturer, The University of Adelaide
Support Vector MachinesMachine LearningAutonomous VehiclesNetwork Traffic Classification