KC-BFPRL: Knowledge-Guided Multi-UAV Collaboration for Grassland Restoration via Bilevel Formerpointer-Based Reinforcement Learning

πŸ“… 2026-08-17
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of maximizing coverage in multi-UAV grassland restoration under payload energy consumption and ecological heterogeneity constraints. We propose a knowledge-guided bilevel reinforcement learning framework that hierarchically decouples task allocation from path planning by integrating Transformers, Pointer Networks, and ecological prior rules to effectively mitigate cold-start issues and ensure constraint satisfaction. Experimental results demonstrate that the proposed framework achieves a 0% optimality gap in complex scenarios while attaining decision-making speeds nearly three times faster than MAPDP. Exhibiting strong robustness and real-time performance, this approach significantly outperforms existing baseline methods, offering an efficient solution for constrained multi-agent coordination in ecological restoration tasks.
πŸ“ Abstract
Multi-unmanned aerial vehicle (UAV) systems provide scalable service platforms for large-scale environmental tasks, such as grassland ecosystem restoration. However, coordinating fleet operations requires solving the restoration area maximization problem (RAMP). This non-linear combinatorial optimization challenge is complicated by payload-dependent energy dynamics and heterogeneous ecological degradation. We propose a novel knowledge-guided collaborative bilevel formerpointer reinforcement learning framework (KC-BFPRL) to address this complexity. Using a hierarchical paradigm, KC-BFPRL decomposes RAMP into global task allocation and local restoration planning, with the latter further divided into upper-level trajectory planning and lower-level restoration area allocation. Our specialized architecture pairs featuring a Transformer-based encoder that fuses static environmental features with dynamic UAV states, and a Pointer Network decoder trained via a robust actor-critic framework. By embedding ecological priority rules and heuristic logic, KC-BFPRL achieves a structured warm-start, solving the RL cold-start problem while ensuring strict constraint satisfaction. Extensive experiments demonstrate that KC-BFPRL consistently outperforms state-of-the-art baselines, achieving superior objective values and efficiency. It maintains a $0.00\%$ optimality gap in the most complex scenarios U8-R160 and operates nearly three times faster than MAPDP, validating its robustness, scalability, and real-time applicability for large-scale automated ecological restoration.
Problem

Research questions and friction points this paper is trying to address.

Restoration Area Maximization Problem
Multi-UAV Collaboration
Grassland Restoration
Non-linear Combinatorial Optimization
Payload-dependent Energy Dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bilevel Reinforcement Learning
Knowledge-Guided RL
Transformer-Pointer Network
Multi-UAV Collaboration
Structured Warm-Start
Dongbin Jiao
Dongbin Jiao
Lanzhou University
Network Resource OptimizationUAV NetworksLow-altitude EconomyEvolutionary Learning
X
Xianyi Wang
School of Information Science and Engineering, Lanzhou University, Lanzhou, 730000, P. R. China
Y
Yuchen Yuan
School of Information Science and Engineering, Lanzhou University, Lanzhou, 730000, P. R. China
Weibo Yang
Weibo Yang
School of Automobile, Chang’an University, Xi’an, 710064, P. R. China
P
Peng Yang
Department of Statistics and Data Science, Southern University of Science and Technology, Shenzhen 518055, P. R. China
P
Peng Zhao
School of Information Science and Engineering, Lanzhou University, Lanzhou, 730000, P. R. China
Zhanhuan Shang
Zhanhuan Shang
State Key Laboratory of Grassland Agro-Ecosystem, College of Ecology, Lanzhou University, Lanzhou, 730000, P. R. China
Shi Yan
Shi Yan
Eindhoven University of Technology
Optical communicationfiber opticsSignal processing