Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究审计了八个网络安全基准在不同语言模型上的表现,揭示了评分依赖于评估流程配置的问题,并提出标准化评估流程以提高模型评价可靠性。
📝 Abstract
Large language model (LLM) benchmarks are often treated as fixed datasets with stable scores, yet their outcomes depend on configurable evaluation pipelines. We audit eight cybersecurity benchmarks across 10 proprietary, open-weight, and cybersecurity-specialized LLMs. By modeling benchmarks as measurement pipelines, we identify 15 systematic failure modes and show that a single pipeline choice can change a model's score by more than 80 percentage points and substantially alter model rankings. At the cross-benchmark level, two semantically similar task pairs rank the same models differently because of incompatible evaluation conventions. Under an evaluation harness that standardizes pipeline choices while preserving task semantics, nine of 10 models shift by at least three ranks on at least one benchmark. These results show that cybersecurity LLM benchmark scores are pipeline-dependent and motivate pipeline-aware auditing as a core requirement for reliable model evaluation.
Problem

Research questions and friction points this paper is trying to address.

cybersecurity
large language model
benchmark
pipeline
evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

pipeline-dependent
cybersecurity LLM benchmarks
systematic failure modes
evaluation harness
🔎 Similar Papers
No similar papers found.