A New Framework for Cybersecurity Refusals in AI Agents
This study addresses the critical gap in current AI agents’ inability to appropriately refuse harmful requests in offensive cybersecurity tasks, where an overemphasis on task completion often overrides safety considerations. We formally define the refusal boundary in this context for the first time, propose evaluable refusal criteria and a taxonomy, and introduce the first evaluation framework specifically designed to assess AI refusal behavior in offensive security scenarios. Leveraging large language model–based agent architectures, we conduct adversarial testing and robustness evaluations across diverse cyberattack settings on eight state-of-the-art models. Our findings reveal that only GPT-5.2 and GPT-5.1 Codex demonstrate meaningful refusal capabilities, while the remaining six models exhibit virtually no refusal behavior, underscoring a severe deficiency in current models’ safety alignment for offensive cybersecurity applications.