Fairis: Fairness-Aware Aggregation with Provable Influence Containment against Fairness Poisoning Attacks in Collaborative Machine Learning
This work addresses fairness poisoning attacks in collaborative machine learning, where malicious clients degrade group fairness while preserving model accuracy to evade accuracy-based defenses. To counter this threat, the authors propose Fairis, a server-side dynamic reweighting mechanism that adjusts client update weights based on their local fairness metrics—normalized using Equal Opportunity Difference—and integrates norm clipping with a safety parameter η for robust control. Fairis provides the first provable bound on the impact of fairness poisoning attacks and offers three theoretical guarantees: monotonic weight decay, demographic participation, and non-manipulability. It also resists collusion by a minority of adversarial clients. Experiments on the Taiwanese Credit dataset demonstrate that Fairis reduces the influence of stealthy attackers by 41%–54% while consistently assigning positive weights to honest clients, substantially outperforming existing approaches.