JavaVulBench: A Java Vulnerability Benchmark with Realistic Splits, a Unified Multi-Backend Harness, and a Leakage-Aware Evaluation Mode

๐Ÿ“… 2026-07-02
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the critical limitations of existing Java vulnerability detection benchmarks, which inadequately balance realistic scenario partitioning, evaluation consistency, and data leakage prevention. To bridge this gap, we construct a high-quality dataset comprising approximately 30,600 methods spanning 1,740 CVEs, introduce a leakage-aware evaluation protocol alongside a pretraining contamination auditing mechanism, and design five practical data splitting strategies. We further establish the first unified evaluation framework that seamlessly integrates encoder-only models, local generative models, and API-based large language models, enabling single-command, multi-backend cross-model assessment for CodeBERT, GraphCodeBERT, UniXcoder, DeepSeek-Coder, and major LLM APIs. The project releases all data, code, fine-tuned models, and twelve reference detectors, accompanied by a complete evaluation pipeline demonstration.
๐Ÿ“ Abstract
We release \textsc{JavaVulBench}, a benchmark dataset and evaluation harness for Java vulnerability detection. The dataset contains $\sim$30{,}600 Java methods spanning 1{,}740 CVEs and 700+ projects, labelled at both method and line granularity, with per-CVE publication dates and five realistic split strategies: random, project-disjoint, temporal, deduplicated, and unseen CWE-family. The harness provides a single \texttt{LlmPrediction} schema across three backend families (encoder classifiers, local generative models served by Ollama, and API-served LLMs routed through OpenRouter) so that twelve reference detectors CodeBERT, GraphCodeBERT, UniXcoder, DeepSeek-Coder-1.3B, and eight API/open-weight LLMs (GPT-4o, GPT-4.1-mini, Claude Sonnet~4, DeepSeek-v3, DeepSeek-Coder-v2, Qwen-2.5-Coder-14B/7B, CodeLlama-13B) are evaluated under identical conditions from a single command. A pre-training contamination audit is shipped alongside every model so users can separate genuinely unseen test CVEs from potentially memorised ones. Data, code, and fine-tuned checkpoints are archived on Zenodo [31] and short demonstration video is available on YouTube (https://www.youtube.com/watch?v=nMTX\_hqkuoM) https://www.youtube.com/watch?v=nMTX_hqkuoM.
Problem

Research questions and friction points this paper is trying to address.

Java vulnerability detection
benchmark dataset
realistic data splits
evaluation leakage
LLM evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

JavaVulBench
realistic data splits
unified multi-backend harness
leakage-aware evaluation
vulnerability detection benchmark
๐Ÿ’ผ Related Jobs
No related jobs found.
N
Norbert Sรกndor Szolnoki
University Of Szeged, Szeged, Hungary
G
Gรกbor Antal
University Of Szeged, Szeged, Hungary