🤖 AI Summary
In federated learning (FL), frequent client participation in model updates leads to cumulative leakage of sensitive information, rendering systems vulnerable to membership/attribute inference and model inversion attacks. To address this, we propose a quantum-inspired QUBO (Quadratic Unconstrained Binary Optimization)-based dynamic client selection mechanism—the first to formulate FL client selection as a binary optimization problem. Our method employs a validation-set-driven, fine-grained update screening strategy to suppress participation frequency of high-risk clients, thereby achieving privacy-utility co-optimization without compromising global model accuracy. Experiments on MNIST (300 clients) and CINIC-10 (30 clients) demonstrate that our approach reduces per-round privacy exposure by 95.2% (49% cumulatively) while maintaining or even exceeding the accuracy of full aggregation—despite rejecting 147 updates. On CINIC-10, per-round and cumulative privacy gains reach 82% and 33%, respectively.
📝 Abstract
Federated learning (FL) is a widely used method for training machine learning (ML) models in a scalable way while preserving privacy (i.e., without centralizing raw data). Prior research shows that the risk of exposing sensitive data increases cumulatively as the number of iterations where a client's updates are included in the aggregated model increase. Attackers can launch membership inference attacks (MIA; deciding whether a sample or client participated), property inference attacks (PIA; inferring attributes of a client's data), and model inversion attacks (MI; reconstructing inputs), thereby inferring client-specific attributes and, in some cases, reconstructing inputs. In this paper, we mitigate risk by substantially reducing per client exposure using a quantum computing-inspired quadratic unconstrained binary optimization (QUBO) formulation that selects a small subset of client updates most relevant for each training round. In this work, we focus on two threat vectors: (i) information leakage by clients during training and (ii) adversaries who can query or obtain the global model. We assume a trusted central server and do not model server compromise. This method also assumes that the server has access to a validation/test set with global data distribution. Experiments on the MNIST dataset with 300 clients in 20 rounds showed a 95.2% per-round and 49% cumulative privacy exposure reduction, with 147 clients'updates never being used during training while maintaining in general the full-aggregation accuracy or even better. The method proved to be efficient at lower scale and more complex model as well. A CINIC-10 dataset-based experiment with 30 clients resulted in 82% per-round privacy improvement and 33% cumulative privacy.