Revisiting Vul-RAG: Reproducibility and Replicability of RAG-based Vulnerability Detection with Open-Weight Models
This study addresses the limited reproducibility and generalizability of existing large language model (LLM)-based vulnerability detection approaches, which often rely on closed-source models and proprietary APIs. The authors systematically reproduce the Vul-RAG framework in a fully local, open-weight setting and, for the first time, evaluate its feasibility across a diverse set of open-source LLMs—including code-specific, general-purpose, and reasoning-oriented models—using a standardized evaluation protocol. Experimental results reveal that all models converge to a performance plateau around 0.30 pairwise accuracy, indicating a saturation effect between model scale and vulnerability detection efficacy. This challenges the prevailing “bigger is better” assumption and suggests that merely increasing model capacity yields diminishing returns in improving detection performance.