Agentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration Testing

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过评估十种常用安全工具,系统化分析了基于大语言模型的代理在渗透测试中的失败模式,并提出了解决策略。
📝 Abstract
Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools. As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures. We systematize these failures through a hands-on evaluation of ten widely used static, dynamic, cloud, orchestration, and AI red-teaming tools for unattended pipelines. We introduce a four-dimensional Integration Friction Index that separates one-time engineering cost from recurring organisational, legal, and maintenance cost. We then derive quantitative regularities that explain recurring failure modes. Modelling an agentic security system as stochastic LLM policies wrapped by a deterministic mediator, we show that long-lived sessions lose resident evidence with phase count, while short-lived sub-agents extend the usable horizon according to the compression ratio between raw evidence and its summary. We show that a two-stage verdict cascade multiplies scorer likelihood ratios, but provides little benefit when scorer errors correlate. We show that treating unevaluable outcomes as attack failures biases downstream measurements toward evasive and severe responses. We formulate planner-versus-worker model routing as a knapsack problem and derive a closed-form execution cap for heavy-tailed tools, eta* = alpha v/c. Finally, we show why scope and budget enforcement cannot be delegated to system prompts: prompts do not constrain what actually executes. Inspectra, our implemented platform, serves as a worked instantiation, with mechanisms labelled shipped, partial, or planned, including those that did not work.
Problem

Research questions and friction points this paper is trying to address.

agentic security
large-language-model
penetration testing
operational failures
security tools
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Security
Integration Friction Index
Stochastic LLM Policies
Verdict Cascade
Knapsack Problem
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
I
Israt Moyeen Noumi
Department of Computer Science and Engineering, Ahsanullah University of Science and Technology, Dhaka, Bangladesh
T
Tarannum Ahmed Nowshin
Department of Computer Science and Engineering, BRAC University, Dhaka, Bangladesh
M
Md. Mehedi Hasan Nipu
Department of Electrical and Computer Engineering, North South University, Dhaka, Bangladesh
M
Mohammad Sakib Mahmood
Department of Computer Science, Missouri State University, Springfield, MO 65897 USA
M
Md. Jakir Hossain
Center for Advanced Analytics, COE for Artificial Intelligence, Faculty of Engineering and Technology, Multimedia University, Melaka 75450, Malaysia
M
M. F. Mridha
Department of Computer Science, American International University-Bangladesh, Dhaka, Bangladesh