SoK: Privacy Attacks on Machine Learning via Explainable AI

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了通过可解释AI对机器学习模型的隐私攻击问题,系统分析了25项研究,并提出了五种攻击路径及相应的防御策略。
📝 Abstract
Machine learning explanations reveal model behavior beyond predictions, creating attack surfaces for model confidentiality and data privacy. We systematize 25 studies that exploit explanations for model extraction, membership inference, and model inversion, treating attribute inference as partial inversion. Existing work is often labeled only black- or white-box, obscuring substantial differences in what explanation signal reaches an adversary. We therefore separate model knowledge from explanation acquisition and identify five paths: target-released, attacker-derived, secondary disclosure, privileged access, and released global artifacts. Across these paths, explanations reduce extraction cost, expose membership signals through explanation statistics, recourse distance, and explanation-guided robustness, and support spatial or algebraic reconstruction of private inputs. We compare system and threat models, explanation signals, auxiliary knowledge, target models, modalities, query budgets, evaluation metrics, reported performance, and defenses. Our analysis shows that no explanation family is uniformly unsafe and no defense is uniformly effective. Risk depends on which signal is exposed, how it is acquired, which asset is targeted, and what the attacker already knows. We argue that explanation privacy should therefore be evaluated as an end-to-end disclosure problem, with defenses matched to the acquisition path and protected asset.
Problem

Research questions and friction points this paper is trying to address.

Privacy Attacks
Machine Learning
Explainable AI
Model Confidentiality
Data Privacy
Innovation

Methods, ideas, or system contributions that make the work stand out.

model extraction
membership inference
explanation paths
privacy attacks
end-to-end disclosure