ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the vulnerability of large language model (LLM) agents to indirect prompt injection attacks embedded in environmental states during external tool use, a risk inadequately captured by existing evaluation methods due to limited scalability and diversity. To bridge this gap, we propose ToolHazard—the first scalable framework for synthesizing adversarial environments—leveraging a coordinated triad of an environment simulator, an attacker agent, and a user simulator to automatically generate stateful, executable attack scenarios. ToolHazard autonomously identifies injection points and crafts targeted payloads, enabling security evaluation in complex, long-horizon tasks without manual intervention. It supports cross-domain seed expansion and produces safety-aligned training data. Experiments reveal significant vulnerabilities in current agents, with attack timing and location critically influencing success rates. Models fine-tuned on ToolHazard-generated data demonstrate markedly improved robustness on ToolHazard-Bench and AgentDojo while preserving standard task performance.
📝 Abstract
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering and supports expansion with additional seed domains and compute. Through an Environment Simulator, an Attacker Agent, and a User Simulator, ToolHazard synthesizes executable stateful environments, discovers viable injection points and generates environment-specific payloads, and constructs state-grounded long-horizon tasks. Based on ToolHazard, we build **ToolHazard-Bench** for stress-testing agents under complex workflows and diverse environmental attacks. Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness. Moreover, ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
indirect prompt injection
adversarial environments
security evaluation
tool integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

adversarial environment synthesis
indirect prompt injection
LLM-based agents
scalable security evaluation
alignment data generation
Yutao Mou
Yutao Mou
Peking University
AI SafetyLLM Alignment
Pengfei Yang
Pengfei Yang
Institute of Software, Chinese Academy of Sciences
Probabilistic model checkingDNN verification
Z
Zhe Yin
Beijing University of Posts and Telecommunications
Z
Zhangchi Xue
National Engineering Research Center for Software Engineering, Peking University
X
Xiaotian Luan
Weixin AI, Tencent Inc.
D
Dingyao Yu
National Engineering Research Center for Software Engineering, Peking University
T
Tong Zhang
Weixin AI, Tencent Inc.
Shikun Zhang
Shikun Zhang
北京大学
Wei Ye
Wei Ye
Peking University
Software EngineeringNatural Language Processing