Hierarchical Reinforcement Learning for Next Generation of Multi-AP Coordinated Spatial Reuse
This work addresses the challenge in multi-access point (AP) coordinated spatial reuse (C-SR), where high coordination overhead and complex parameter optimization hinder the simultaneous achievement of high throughput and fairness. To this end, the paper proposes a two-layer multi-armed bandit (MAB) algorithm that, for the first time, introduces hierarchical reinforcement learning into the multi-AP C-SR setting. The proposed framework jointly optimizes scheduling, power control, and link adaptation through a hierarchical structure, significantly reducing coordination overhead while meeting diverse quality-of-service (QoS) requirements. System-level simulations demonstrate that the scheme not only enhances aggregate network throughput but also improves fairness in resource allocation among users, offering a robust and efficient solution for next-generation Wi-Fi systems.