Pareto-Aware Hierarchical Reinforcement Learning for Online Resource Allocation in RIS-assisted Large-Scale IoT Systems

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the computational complexity of real-time resource allocation in reconfigurable intelligent surface (RIS)-assisted massive Internet-of-Things systems, which arises from high-dimensional beamforming and the non-convex Pareto-optimal rate region. To tackle this challenge, the paper proposes a dimensionality-reduced hierarchical reinforcement learning framework, termed PAAERL. The approach uniquely embeds the Pareto front into low-dimensional weight vectors to preserve geometric information and leverages an autoencoder to achieve a second-stage compression from a priority space to a continuous latent action space. By integrating Pareto-aware reinforcement learning with hierarchical optimization, the method substantially reduces computational overhead, significantly shortens offline training time, and accelerates online policy convergence. Evaluated in multi-user mobile edge computing scenarios, PAAERL outperforms state-of-the-art methods, demonstrating enhanced system scalability and resource allocation efficiency.
📝 Abstract
With the rapid evolution of 5G and emerging 6G networks, reconfigurable intelligent surfaces (RIS) have become a critical technology for enhancing wireless communication scenarios. However, optimizing RIS-assisted multi-user systems typically introduces high-dimensional physical layer variables and non-convex Pareto-optimal rate sets, posing severe computational challenges for real-time applications. To address these limitations, this paper proposes a dimension-reduced, hierarchical reinforcement learning (RL) framework, termed Pareto-aware autoencoder-assisted RL (PAAERL), to optimize online resource allocation in RIS-assisted Internet of Things (IoT) networks. Our approach first substitutes high-dimensional continuous RIS beamforming variables with lower-dimensional weight vectors that strictly represent the Pareto-optimal frontier, theoretically avoiding geometric information loss across both convex and non-convex rate regions. To further mitigate the curse of dimensionality in dense networks, an autoencoder architecture is integrated to execute a secondary, data-driven compression phase, mapping the priority space into a highly condensed continuous latent action space. Extensive simulations conducted across practical communication scenarios, including multi-user mobile edge computing (MEC) networks, demonstrate that the proposed PAAERL framework drastically reduces offline training times, accelerates online policy convergence, and significantly decreases overall network costs compared to state-of-the-art benchmarks, underscoring its exceptional scalability and practical viability for next-generation intelligent IoT environments.
Problem

Research questions and friction points this paper is trying to address.

Reconfigurable Intelligent Surfaces
Resource Allocation
Pareto Optimality
Large-Scale IoT
Online Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pareto-aware optimization
hierarchical reinforcement learning
reconfigurable intelligent surfaces (RIS)
autoencoder-based dimensionality reduction
online resource allocation
W
Wenhan Xu
Internet of Things Thrust, Hong Kong University of Science and Technology (Guangzhou), Guangzhou, Guangdong 511400, China
Jiashuo Jiang
Jiashuo Jiang
Hong Kong University of Science and Technology
operations researchoperations managementoptimizationapproximation algorithmsmachine learning
D
Danny H. K. Tsang
Internet of Things Thrust, Hong Kong University of Science and Technology (Guangzhou), Guangzhou, Guangdong 511400, China, and also with the Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology, Clear Water Bay, Hong Kong SAR, China