Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入上下文分割方法,解决了本地部署的小型语言模型在处理网络安全CTF任务时因上下文膨胀导致的问题。
📝 Abstract
The proliferation of highly capable open-weight Small Language Models (SLMs) democratizes access to advanced cybersecurity capabilities, posing a escalating risk as these models can bypass proprietary API guardrails when deployed locally. However, SLMs deployed as autonomous agents often struggle with long-horizon, exploratory tasks like cybersecurity Capture The Flag (CTF) challenges due to context bloat and cognitive degradation from accumulated tool-call outputs. To understand and mitigate this cybersecurity threat, we introduce \textit{context segmentation}, a two-level agentic framework that divides complex exploitation tasks into manageable, contextually isolated sub-problems. Evaluating on the \texttt{picoCTF} dataset using memory-constrained \texttt{gemma-4} models, we demonstrate that for the E4B model, our strategy acts as an intelligent search, achieving competitive rewards with superior token efficiency compared to brute-force retries, and successfully solving 18.52\% of tasks that standard agentic execution fails to complete. Code is available at https://github.com/9xeb/context-segmentation.
Problem

Research questions and friction points this paper is trying to address.

Small Language Models
cybersecurity
context bloat
cognitive degradation
autonomous agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

context segmentation
small language models
cybersecurity CTF
gemma-4
autonomous agents
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sebastiano Nordio
Independent
M
Michele Lotto
University of Genoa, Via All'Opera Pia, 16145 Genoa, Italy