Logo image
BenchGuard: A Two-Layer Framework for Contamination-Aware Evaluation of LLM Vulnerability Detection
Conference proceeding

BenchGuard: A Two-Layer Framework for Contamination-Aware Evaluation of LLM Vulnerability Detection

Adiba Mahmud, Yasmeen Rawajfih, Ross Arnold and Hossain Shahriar
Data and Applications Security and Privacy XL, pp.408-420
Lecture Notes in Computer Science
IFIP WG 11.3 Annual Conference, DBSec (Arlington, Virginia, USA, 07/28/2026–07/30/2026)
07/21/2026

Metrics

1 Record Views

Abstract

Benchmark integrity Data contamination Data governance LLM evaluation Vulnerability classification Cybersecurity Machine Learning
Security benchmarks built from publicly available vulnerability databases are central to how large language models are evaluated in the cybersecurity domain, yet the validity of those evaluations rests on an assumption that is rarely examined: that evaluation data is meaningfully distinct from what models encountered during pre-training. When this assumption fails, benchmark scores reflect memorization rather than reasoning, and the conclusions drawn from them may not transfer to practice. We present BenchGuard, a two-layer framework that makes this problem measurable and correctable. The first layer governs training data quality through contamination-aware weighting, combining near-duplicate detection, boilerplate filtering, and similarity scoring. The second layer applies formal statistical tests to evaluation outputs to identify results inconsistent with genuine reasoning at observed accuracy levels. Across five coordinated experiments, we confirm that training-evaluation overlap causes genuine, quantified performance inflation that disappears on non-overlapping records and is therefore attributable to leakage rather than generalizable model improvement. A controlled injection experiment further decomposes this inflation into two separable mechanisms. A cross-architecture comparison shows that a classical pipeline operating under proper data governance outperforms a five-encoder neural ensemble by over thirteen percentage points, a result that challenges the assumption that architectural complexity is the primary driver of evaluation performance. BenchGuard is offered as a reusable methodology for any setting where training and evaluation data are drawn from the same public vulnerability corpus.

Details

Logo image