🤖 AI Summary
This study addresses the lack of systematic alignment between Deep Reinforcement Learning (DRL) methodologies and practical deployment in 6G Open AI-RAN. We propose an O-RAN-aware framework that integrates multi-agent collaboration, federated learning, and Sim-to-Real transfer techniques to systematically review wireless resource management and coordination scenarios. As the first comprehensive survey dedicated to this domain, this work establishes a standardized problem modeling system and taxonomy. Furthermore, it identifies critical research directions, including sample efficiency, safety and trustworthiness, and scalable control. Ultimately, this paper provides a systematic roadmap for both theoretical research and engineering implementation of intelligent RAN in 6G networks, bridging the gap between academic DRL innovations and real-world O-RAN deployment requirements.
📝 Abstract
The evolution toward sixth-generation (6G) networks is transforming the radio access network (RAN) into a programmable and intelligent control platform that must continuously adapt to heterogeneous services, dynamic environments, and competing performance objectives. Open Radio Access Network (O-RAN) provides the open interfaces, disaggregated architecture, and multi-timescale control loops needed to support this transformation, while deep reinforcement learning (DRL) offers a natural framework for optimizing sequential decisions under uncertainty. However, existing surveys either address artificial intelligence (AI) and machine learning (ML) in O-RAN broadly or focus on isolated DRL use cases, leaving a gap in the systematic connection between DRL methodology, O-RAN architecture, and operational deployment. To the best of our knowledge, this article presents the first dedicated and comprehensive survey of DRL for Open AI-RAN. We review the foundations of model-free, model-based, offline, safe, multi-agent, federated, and transfer learning, and provide an O-RAN-aware framework for formulating RAN control problems through states, observations, actions, rewards, constraints, and temporal structure. We classify DRL applications across radio resource management, mobility management, interference control, traffic steering, energy efficiency, network slicing, integrated sensing and communication, security, and massive MIMO. We further examine multi-agent and federated coordination, foundation models and agentic AI, trustworthy DRL, sim-to-real transfer, continual adaptation, resource-efficient inference, and reinforcement learning operations. Finally, we review experimental platforms, benchmarks, standards, and industry activities, and identify research directions toward sample-efficient, safe, scalable, interoperable, and deployable DRL control for 6G Open AI-RAN.