How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过引入CTF-ABACUS框架,解决了评估语言模型代理在网络安全攻防演练中实际能力的问题,此框架能够追踪和审计代理的行为路径。
📝 Abstract
Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of autonomous language-model agents. Evaluations rely on shallow binary judgments or aggregate scores, overlooking the agent's trajectory to the flag. Consequently actual exploitation is conflated with direct flag exposure, memorized recall, external lookup, guessing, and unsupported claims, potentially overstating the agent's cybersecurity capability. We introduce CTF-ABACUS, a trace-based agent auditing framework that reconstructs each run as an evidence-grounded solve profile. By decomposing agent actions into penetration-testing phases and categorical techniques, it identifies where exploitation occurs, where the flag first appears, and whether the recovered flag is supported by demonstrated behavior. Aggregating solve profiles across agents yields challenge signatures that reveal whether success was achieved via the intended exploit or via shortcut pathways. We apply CTF-ABACUS to 1,435 CTF attempts by six frontier and open-source models on 240 challenges, yielding 2,870 solve profiles under two judge lenses. Trace-verified exploits account for only 62-87% of recovered flags across benchmarks, while shortcut recoveries follow substantially shallower trajectories. These findings shift CTF evaluation from counting recovered flags to verifying demonstrated exploitation and provide a basis for designing benchmarks that better isolate the offensive capabilities.
Problem

Research questions and friction points this paper is trying to address.

Capture-the-Flag
offensive security
autonomous language-model agents
exploitation
benchmark evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

trace-based agent auditing
penetration-testing phases
categorical techniques
challenge signatures
trace-verified exploits
🔎 Similar Papers
K
Kimberly Milner
NYU Tandon School of Engineering
M
Minghao Shao
NYU Tandon School of Engineering, NYU Abu Dhabi
N
Nanda Rani
CISPA - Helmholtz Center for Information Security, Saarbrücken
H
Haoran Xi
NYU Tandon School of Engineering
V
Venkata Sai Charan Putrevu
Indian Institute of Technology Tirupati
M
Meet Udeshi
NYU Tandon School of Engineering
S
Sandeep K. Shukla
International Institute of Information Technology Hyderabad
Prashanth Krishnamurthy
Prashanth Krishnamurthy
Research Scientist, New York University
roboticscontrol systemscyber-physical systems
Farshad Khorrami
Farshad Khorrami
Professor of Electrical and Computer Engineering, NYU
RoboticsControl SystemsCyber Physical System SecurityDecentralized Control
Muhammad Shafique
Muhammad Shafique
Professor, ECE, New York University (AD-UAE, Tandon-USA), Director eBRAIN Lab
Embedded Machine LearningBrain-Inspired ComputingRobust & Energy-Efficient System DesignSmart
R
Ramesh Karri
NYU Tandon School of Engineering