Two Truths and A Lie? Benchmarking Off-the-Shelf LLMs for Requirements Quality Assessment: Performance, False Alarms, and Misses

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究评估了现成的大型语言模型在需求质量评估中的表现,发现这些模型错误率高且性能不稳定,不适合独立进行需求质量评估。
📝 Abstract
Requirements engineering (RE) governs the quality of everything downstream in systems engineering (SE); defective requirements that survive review cycles propagate into design rework, schedule delays, and cost overruns. Because requirements are often written in natural language, recent advances in generative AI have raised expectations that large language models (LLMs) can absorb requirement quality assessment, a task otherwise slow and human expertise-intensive. Yet empirical evidence on whether LLMs can be trusted to do so remains scarce. This study presents the first benchmarking analysis of off-the-shelf LLM performance for requirement quality evaluation. Against an expert-derived ground truth built on INCOSE quality criteria, we evaluate ten models spanning two families (OpenAI and Anthropic) and five generations each, across one hundred independent runs, two requirement sets, and five sampling temperatures. Four contributions follow. First, we quantify a strongly asymmetric error profile: across all models and runs, the best-performing Anthropic model detects a median of only 47% of expert-identified issues while false-flagging 11%. Second, performance degrades significantly where SE judgment is required, as necessity and correctness issues are almost always missed. Third, generational progress is non-monotonic, so newer models cannot be assumed better. Fourth, this error behavior shifts only modestly and non-monotonically across sampling temperatures, indicating characteristic model deficiencies rather than inherent stochasticity. Off-the-shelf LLMs are therefore not yet trustworthy autonomous evaluators. Findings also warrant caution for Agentic AI developers: orchestrating these LLM modules in specialized architectures risks compounding these deficiencies rather than correcting them. Their defensible near-term role is human-in-the-loop decision support.
Problem

Research questions and friction points this paper is trying to address.

Requirements Quality Assessment
Large Language Models
False Alarms
Misses
Human Expertise
Innovation

Methods, ideas, or system contributions that make the work stand out.

benchmarking
requirement quality assessment
large language models
asymmetric error profile
human-in-the-loop
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jannatul Shefa
Grado Department of Industrial and Systems Engineering, Virginia Tech, Blacksburg, Virginia, USA
A
Alejandro Salado
Department of Systems and Industrial Engineering, University of Arizona, Tucson, Arizona, USA
Paul Wach
Paul Wach
Department of Systems and Industrial Engineering, University of Arizona, Tucson, Arizona, USA
Taylan G. Topcu
Taylan G. Topcu
Assistant Professor of Systems Engineering & Analysis @ Virginia Tech, the Grado Department of ISE
Systems EngineeringSociotechnical SystemsDigital EngineeringModularity