Neutral Agent-based Adversarial Policy Learning against Deep Reinforcement Learning in Multi-party Open Systems

πŸ“… 2025-10-12
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
In multi-agent open systems, deep reinforcement learning (DRL) agents are vulnerable to adversarial attacks; however, existing methods require either full environmental control or direct interaction with the victim agent, severely limiting practical applicability. Method: This paper proposes an indirect adversarial strategy learning framework leveraging a *neutral agent*β€”an autonomous entity deployed in the shared environment that neither interacts with the victim nor controls the environment, yet exerts implicit interference to mislead the victim’s decision-making. The approach integrates DRL with adversarial policy optimization to generate robust, generalizable deceptive behaviors. Results: Experiments on SMAC and Highway-env demonstrate cross-scenario effectiveness: the method consistently misleads high-performance victim agents while significantly improving attack stealthiness, transferability, and real-world feasibility compared to conventional approaches reliant on strong assumptions.

Technology Category

Application Category

πŸ“ Abstract
Reinforcement learning (RL) has been an important machine learning paradigm for solving long-horizon sequential decision-making problems under uncertainty. By integrating deep neural networks (DNNs) into the RL framework, deep reinforcement learning (DRL) has emerged, which achieved significant success in various domains. However, the integration of DNNs also makes it vulnerable to adversarial attacks. Existing adversarial attack techniques mainly focus on either directly manipulating the environment with which a victim agent interacts or deploying an adversarial agent that interacts with the victim agent to induce abnormal behaviors. While these techniques achieve promising results, their adoption in multi-party open systems remains limited due to two major reasons: impractical assumption of full control over the environment and dependent on interactions with victim agents. To enable adversarial attacks in multi-party open systems, in this paper, we redesigned an adversarial policy learning approach that can mislead well-trained victim agents without requiring direct interactions with these agents or full control over their environments. Particularly, we propose a neutral agent-based approach across various task scenarios in multi-party open systems. While the neutral agents seemingly are detached from the victim agents, indirectly influence them through the shared environment. We evaluate our proposed method on the SMAC platform based on Starcraft II and the autonomous driving simulation platform Highway-env. The experimental results demonstrate that our method can launch general and effective adversarial attacks in multi-party open systems.
Problem

Research questions and friction points this paper is trying to address.

Developing adversarial attacks without direct victim interaction
Overcoming impractical full environment control assumptions
Enabling attacks in multi-party open systems indirectly
Innovation

Methods, ideas, or system contributions that make the work stand out.

Neutral agent indirectly influences victim agents
Adversarial policy without direct victim interaction
Attacks through shared environment in open systems
πŸ”Ž Similar Papers
Q
Qizhou Peng
State Key Laboratory of Cyberspace Security Defense, Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Y
Yang Zheng
State Key Laboratory of Cyberspace Security Defense, Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Y
Yu Wen
State Key Laboratory of Cyberspace Security Defense, Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Y
Yanna Wu
State Key Laboratory of Cyberspace Security Defense, Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Y
Yingying Du
State Key Laboratory of Cyberspace Security Defense, Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China