π€ AI Summary
This study addresses the challenge of maximizing coverage in multi-UAV grassland restoration under payload energy consumption and ecological heterogeneity constraints. We propose a knowledge-guided bilevel reinforcement learning framework that hierarchically decouples task allocation from path planning by integrating Transformers, Pointer Networks, and ecological prior rules to effectively mitigate cold-start issues and ensure constraint satisfaction. Experimental results demonstrate that the proposed framework achieves a 0% optimality gap in complex scenarios while attaining decision-making speeds nearly three times faster than MAPDP. Exhibiting strong robustness and real-time performance, this approach significantly outperforms existing baseline methods, offering an efficient solution for constrained multi-agent coordination in ecological restoration tasks.
π Abstract
Multi-unmanned aerial vehicle (UAV) systems provide scalable service platforms for large-scale environmental tasks, such as grassland ecosystem restoration. However, coordinating fleet operations requires solving the restoration area maximization problem (RAMP). This non-linear combinatorial optimization challenge is complicated by payload-dependent energy dynamics and heterogeneous ecological degradation. We propose a novel knowledge-guided collaborative bilevel formerpointer reinforcement learning framework (KC-BFPRL) to address this complexity. Using a hierarchical paradigm, KC-BFPRL decomposes RAMP into global task allocation and local restoration planning, with the latter further divided into upper-level trajectory planning and lower-level restoration area allocation. Our specialized architecture pairs featuring a Transformer-based encoder that fuses static environmental features with dynamic UAV states, and a Pointer Network decoder trained via a robust actor-critic framework. By embedding ecological priority rules and heuristic logic, KC-BFPRL achieves a structured warm-start, solving the RL cold-start problem while ensuring strict constraint satisfaction. Extensive experiments demonstrate that KC-BFPRL consistently outperforms state-of-the-art baselines, achieving superior objective values and efficiency. It maintains a $0.00\%$ optimality gap in the most complex scenarios U8-R160 and operates nearly three times faster than MAPDP, validating its robustness, scalability, and real-time applicability for large-scale automated ecological restoration.