Logo image
Trust Score Reliability in LLM Security Pipelines: A Severity-Stratified Calibration Study with KEV Validation
Conference proceeding   Peer reviewed

Trust Score Reliability in LLM Security Pipelines: A Severity-Stratified Calibration Study with KEV Validation

Adiba Mahmud, Yasmeen Rawajfih, Ross Arnold and Hossain Shahriar
Proceedings : annual International Computer Software and Applications Conference, pp.3242-3247
Annual Computers, Software, and Applications Conference (COMPSAC), 50th (Madrid, Spain, 07/07/2026–07/10/2026)
08/2026

Metrics

1 Record Views

Abstract

CISA KEV CVSS DevSecOps ensemble uncertainty LLM calibration security pipeline trust calibration vulnerability triage Cybersecurity
Large language models (LLMs) are increasingly integrated into DevSecOps pipelines for vulnerability triage, yet the reliability of their confidence signals across the vulnerability severity spectrum has received limited empirical study. This paper presents a systematic evaluation of a three-stage LLM ensemble pipeline on 500 CVEs drawn from the National Vulnerability Database, using CVSS-derived severity and exploitability labels alongside CISA Known Exploited Vulnerabilities (KEV) membership as external ground truth. We introduce a reproducible, annotation-free ground truth construction methodology that enables large-scale pipeline evaluation without manual labeling. Our central finding is that ensemble trust score reliability is severity-contingent: for Critical-severity CVEs, trust score predicts action accuracy monotonically, reaching \mathbf{1 0 0} \boldsymbol{\%} at threshold \boldsymbol{\geq} \mathbf{0. 8 5} on 46% of the Critical subset. For High-severity CVEs, the relationship is inverted, with \mathbf{6. 7 \%} accuracy at trust \geq \mathbf{0. 9 0} despite strong inter-model agreement. Trust score is uninformative for Medium and Low severity. Validation against 28 KEV CVEs yields 92.9% Stage 3 accuracy, versus 24.8% on non-KEV CVEs, providing a KEV-validated benchmark for LLM-based security triage. These findings challenge the assumption that a single confidence threshold can safely govern autonomous decisions across all severity levels and motivate severity-aware routing.

Details

Logo image