Communication-Aware Multi-Agent Reinforcement Learning for Decentralized Cooperative UAV Deployment
This work addresses the challenge of cooperative deployment for drone swarms under partial observability and intermittent communication by proposing a graph neural network–based multi-agent reinforcement learning approach grounded in the centralized training with decentralized execution (CTDE) paradigm. The method employs a distance-constrained communication graph and introduces agent-entity attention alongside neighbor self-attention mechanisms, enabling efficient coordination using only local observations and messages from nearby agents. This architecture supports zero-shot generalization to formations of varying scales. Experimental results demonstrate that, in the DroneConnect task, a team of five drones achieves 74% area coverage—approaching the offline upper bound provided by mixed-integer linear programming—and significantly outperforms non-communicating baselines in the DroneCombat task, thereby validating the approach’s effectiveness and scalability.