Making Teams and Influencing Agents: Efficiently Coordinating Decision Trees for Interpretable Multi-Agent Reinforcement Learning

📅 2025-05-25

📈 Citations: 0

✨ Influential: 0

career value

239K/year

🤖 AI Summary

Poor interpretability of multi-agent reinforcement learning (MARL) poses critical safety and trust bottlenecks for real-world deployment. To address this, we propose HYDRAVIPER—a novel MARL algorithm featuring (i) a team-level expected performance-driven cooperative training mechanism and (ii) an environment interaction budget-aware adaptive allocation strategy, jointly achieving Pareto-optimal trade-offs between performance and computational efficiency. HYDRAVIPER integrates decision-tree-based policy modeling, multi-agent cooperative optimization, budget-aware reinforcement learning, and interpretable policy distillation to construct a high-performance, inherently interpretable surrogate policy model. Evaluated on collaborative benchmark tasks and traffic signal control, HYDRAVIPER achieves state-of-the-art (SOTA) performance while accelerating inference by multiple orders of magnitude. Crucially, it maintains stable Pareto-frontier performance across varying interaction budgets, demonstrating robustness and practical deployability.

Technology Category

Application Category

📝 Abstract

Poor interpretability hinders the practical applicability of multi-agent reinforcement learning (MARL) policies. Deploying interpretable surrogates of uninterpretable policies enhances the safety and verifiability of MARL for real-world applications. However, if these surrogates are to interact directly with the environment within human supervisory frameworks, they must be both performant and computationally efficient. Prior work on interpretable MARL has either sacrificed performance for computational efficiency or computational efficiency for performance. To address this issue, we propose HYDRAVIPER, a decision tree-based interpretable MARL algorithm. HYDRAVIPER coordinates training between agents based on expected team performance, and adaptively allocates budgets for environment interaction to improve computational efficiency. Experiments on standard benchmark environments for multi-agent coordination and traffic signal control show that HYDRAVIPER matches the performance of state-of-the-art methods using a fraction of the runtime, and that it maintains a Pareto frontier of performance for different interaction budgets.

Problem

Research questions and friction points this paper is trying to address.

Improving interpretability in multi-agent reinforcement learning policies

Balancing performance and computational efficiency in MARL surrogates

Coordinating decision trees for efficient multi-agent team performance

Innovation

Methods, ideas, or system contributions that make the work stand out.

Decision tree-based interpretable MARL algorithm

Coordinates training based on team performance

Adaptively allocates interaction budgets efficiently

🔎 Similar Papers

MIXRTs: Toward Interpretable Multi-Agent Reinforcement Learning via Mixing Recurrent Soft Decision Trees