Large language models (LLMs) are increasingly integrated into DevSecOps pipelines for vulnerability triage, yet the reliability of their confidence signals across the vulnerability severity spectrum has received limited empirical study. This paper presents a systematic evaluation of a three-stage LLM ensemble pipeline on 500 CVEs drawn from the National Vulnerability Database, using CVSS-derived severity and exploitability labels alongside CISA Known Exploited Vulnerabilities (KEV) membership as external ground truth. We introduce a reproducible, annotation-free ground truth construction methodology that enables large-scale pipeline evaluation without manual labeling. Our central finding is that ensemble trust score reliability is severity-contingent: for Critical-severity CVEs, trust score predicts action accuracy monotonically, reaching \mathbf{1 0 0} \boldsymbol{\%} at threshold \boldsymbol{\geq} \mathbf{0. 8 5} on 46% of the Critical subset. For High-severity CVEs, the relationship is inverted, with \mathbf{6. 7 \%} accuracy at trust \geq \mathbf{0. 9 0} despite strong inter-model agreement. Trust score is uninformative for Medium and Low severity. Validation against 28 KEV CVEs yields 92.9% Stage 3 accuracy, versus 24.8% on non-KEV CVEs, providing a KEV-validated benchmark for LLM-based security triage. These findings challenge the assumption that a single confidence threshold can safely govern autonomous decisions across all severity levels and motivate severity-aware routing.
Related links
Details
Title
Trust Score Reliability in LLM Security Pipelines: A Severity-Stratified Calibration Study with KEV Validation
Publication Details
Proceedings : annual International Computer Software and Applications Conference, pp.3242-3247