Submodular Policy Learning for Distributed Task Allocation in Open Multi-Agent Systems

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of relaxed inconsistency in distributed task allocation strategies caused by dynamic membership changes in open multi-agent systems. We propose a partitioned multilinear extension framework and the SubMAPL algorithm, which integrate KL mirror policy learning, local marginal gain estimation, and an open policy transfer mechanism to effectively tackle policy gradient learning and adaptability in dynamic environments. Experimental results demonstrate that this approach significantly outperforms existing policy gradient and online learning baselines in multi-agent coverage tasks. Furthermore, we establish a cumulative utility lower bound that accounts for environmental openness, providing both theoretical guarantees and an efficient solution paradigm for dynamic systems.
📝 Abstract
This paper studies policy learning for distributed task allocation in open multi-agent systems, where agents may join and leave in a time-varying fashion, with submodular stage team utilities. At each time, the active agents select actions from local categorical policies such that the feasible joint agent-action pairs form a partition matroid. Standard continuous relaxations of submodular set functions are based on independent Bernoulli sampling, making them inconsistent with agents' policies.To solve this mismatch, we propose the \emph{partition multilinear extension} (PME), a policy-based relaxation whose continuous support matches feasible actions under categorical policies.We prove that the marginal gains of the stage utility provide an unbiased estimator of the gradient of the PME and that maximizing the PME over action distributions is equivalent to maximizing the stage utilities over agent actions, which are critical to devise principled policy gradient.Building on this, we design \emph{SubMAPL}, a centralized-training decentralized-execution KL-mirror policy-learning method that uses local marginal gains as stochastic PME gradients during training. KL-mirror updates preserve categorical feasibility without Euclidean projection.In the case where agents run tabular-softmax policies, we introduce open policy migration and an open-system KL tracking variation to handle agent arrivals and departures. Using dynamic regret analysis, we establish a lower bound on the cumulative utility which accounts for the openness of the environment and for the gap between optimal stage-wise and global utilities. Simulations on multi-agent coverage demonstrate that SubMAPL outperforms policy-gradient and online-learning baselines.
Problem

Research questions and friction points this paper is trying to address.

Distributed Task Allocation
Open Multi-Agent Systems
Submodular Policy Learning
Continuous Relaxation Mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Partition Multilinear Extension
SubMAPL
KL-mirror Policy Learning
Open Multi-Agent Systems
Dynamic Regret Analysis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Jing Liu
Jing Liu
Technical Institute of Physics and Chemistry, Chinese Academy of Sciences & Tsinghua University
Liquid MetalsThermal PhysicsBioheat TransferBiomedical EngineeringSoft Matters
Luca Ballotta
Luca Ballotta
Postdoc at Delft Center for Systems and Control
Multi-agent systemNetwork control systemsResilient distributed controlControl barrier function
Y
Yangyang Yang
School of Mathematics, East China University of Science and Technology, Shanghai 200237, China
F
Fangfei Li
School of Mathematics and the Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education, East China University of Science and Technology, Shanghai 200237, China
Y
Yang Tang
Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education, East China University of Science and Technology, Shanghai 200237, China
Ruggero Carli
Ruggero Carli
Associate Professor at University of Padova
Control Theory