Hybrid Offline-Online Multi-Agent Decision Transformers for Wireless Resource Management

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该论文提出一种基于决策变换器的混合离线-在线多智能体强化学习框架,用于无线资源管理,通过预训练和在线微调提高性能。
📝 Abstract
This paper develops a hybrid offline-online multi-agent reinforcement learning framework based on decision transformers. The policy is first pretrained offline via supervised sequence modeling of trajectories generated by existing policies, providing a safe and sample-efficient initialization. It is then fine-tuned online using a hybrid objective that incorporates critic-guided gradients, enabling performance improvements beyond the offline policy. To facilitate stable offline-to-online transfer and effective multi-agent coordination, the framework incorporates return-weighted sampling, a critic conditioned on neighbors' actions, and neighborhood-correlated exploration. The approach is fully distributed: both training and execution rely only on local observations and limited information exchange among neighboring agents. Evaluations with dynamic traffic arrivals in two settings: (i) joint scheduling and power allocation and (ii) coordinated beamforming, show that the proposed method achieves quality-of-service (QoS) performance comparable to centralized methods. Moreover, when pretrained on lower-quality datasets, online fine-tuning is also observed to surpass the initial offline policy. These results demonstrate a promising learning-based alternative for wireless resource management.
Problem

Research questions and friction points this paper is trying to address.

multi-agent reinforcement learning
wireless resource management
offline-online hybrid
distributed system
Innovation

Methods, ideas, or system contributions that make the work stand out.

hybrid offline-online
multi-agent reinforcement learning
decision transformers
return-weighted sampling
critic-guided gradients
Y
Yiming Zhang
Department of Electrical and Computer Engineering, Northwestern University, Evanston, IL 60208 USA
Kun Yang
Kun Yang
Google, University of Virginia
Reinforcement LearningLarge Language ModelsCode Optimization
Cong Shen
Cong Shen
Associate Professor, University of Virginia
Multi-armed banditsReinforcement learningFederated learningWireless
D
Dongning Guo
Department of Electrical and Computer Engineering, Northwestern University, Evanston, IL 60208 USA