Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过3R-Bench评估了对话上下文对LLM处理网络安全请求的影响,旨在防止恶意使用同时不拒绝合法帮助。
📝 Abstract
Large Language Models (LLMs) can solve complex problems, but their misuse in high-risk domains can lead to severe consequences. Model providers therefore restrict assistance for potentially harmful requests. Refusing all cybersecurity requests would therefore harm legitimate users. Providers need a mechanism to block malicious use without denying legitimate assistance to defenders. Existing cybersecurity-specific datasets evaluate this mechanism, but none considers the conversational context of a request. We introduce 3R-Bench (Refusal, Repetition, and Revision), a benchmark of 150 real-world cybersecurity requests augmented with two adversarial conversational settings, and evaluate eight LLMs on it. Prior assistant behavior strongly changes responses to an unchanged request: among 376 available pairs from a 400-pair panel, compliance rises from 62.0% after refused history to 85.1% after accepted history. The opposite pattern appears under dialogue decomposition. In comparison, compliance falls from 501/800 direct responses to 172/800 after dialogue; among 738 pairs returning model-authored text in both conditions, the decrease is 45.1 points. Failure feedback recovers only a small fraction of this loss.
Problem

Research questions and friction points this paper is trying to address.

cybersecurity
large language models
conversational context
malicious use
legitimate assistance
Innovation

Methods, ideas, or system contributions that make the work stand out.

3R-Bench
conversational context
cybersecurity requests
🔎 Similar Papers
No similar papers found.