🤖 AI Summary
This work addresses the computational complexity of real-time resource allocation in reconfigurable intelligent surface (RIS)-assisted massive Internet-of-Things systems, which arises from high-dimensional beamforming and the non-convex Pareto-optimal rate region. To tackle this challenge, the paper proposes a dimensionality-reduced hierarchical reinforcement learning framework, termed PAAERL. The approach uniquely embeds the Pareto front into low-dimensional weight vectors to preserve geometric information and leverages an autoencoder to achieve a second-stage compression from a priority space to a continuous latent action space. By integrating Pareto-aware reinforcement learning with hierarchical optimization, the method substantially reduces computational overhead, significantly shortens offline training time, and accelerates online policy convergence. Evaluated in multi-user mobile edge computing scenarios, PAAERL outperforms state-of-the-art methods, demonstrating enhanced system scalability and resource allocation efficiency.
📝 Abstract
With the rapid evolution of 5G and emerging 6G networks, reconfigurable intelligent surfaces (RIS) have become a critical technology for enhancing wireless communication scenarios. However, optimizing RIS-assisted multi-user systems typically introduces high-dimensional physical layer variables and non-convex Pareto-optimal rate sets, posing severe computational challenges for real-time applications. To address these limitations, this paper proposes a dimension-reduced, hierarchical reinforcement learning (RL) framework, termed Pareto-aware autoencoder-assisted RL (PAAERL), to optimize online resource allocation in RIS-assisted Internet of Things (IoT) networks. Our approach first substitutes high-dimensional continuous RIS beamforming variables with lower-dimensional weight vectors that strictly represent the Pareto-optimal frontier, theoretically avoiding geometric information loss across both convex and non-convex rate regions. To further mitigate the curse of dimensionality in dense networks, an autoencoder architecture is integrated to execute a secondary, data-driven compression phase, mapping the priority space into a highly condensed continuous latent action space. Extensive simulations conducted across practical communication scenarios, including multi-user mobile edge computing (MEC) networks, demonstrate that the proposed PAAERL framework drastically reduces offline training times, accelerates online policy convergence, and significantly decreases overall network costs compared to state-of-the-art benchmarks, underscoring its exceptional scalability and practical viability for next-generation intelligent IoT environments.