Logo image
Forensic Attribution and Targeted Regression Detection for AI-Generated Security Patches
Conference proceeding   Peer reviewed

Forensic Attribution and Targeted Regression Detection for AI-Generated Security Patches

Yasmeen Rawajfih, Adiba Mahmud, Ross Arnold and Hossain Shahriar
Proceedings : annual International Computer Software and Applications Conference, pp.2891-2896
Annual Computers, Software, and Applications Conference (COMPSAC), 50th (Madrid, Spain, 07/07/2026–07/10/2026)
08/2026

Metrics

1 Record Views

Abstract

AI-generated patches DevSecOps forensic attribution LLM security Pipelines regression detection static analysis vulnerability remediation Artificial Intelligence or Cybernetics Computer Applications Cybersecurity
As large language models enter security patching workflows, a practical challenge has emerged that existing tooling does not fully address: when an AI-generated patch introduces a new vulnerability, analysts have no way to determine which tool produced it or which category of regression to prioritize. This paper investigates whether AI patching tools leave detectable forensic signatures in the regressions they introduce, and whether those signatures can be used to reduce the static analysis workload without requiring access to the generating model. Working with 1,437 patches from three frontier models across 479 C/C++ vulnerabilities drawn from real CVE records, we find that static patch features enable attribution of a patch to its generating model at accuracy substantially above the uninformed baseline, and that the regression rule sets of all three models converge enough to support a compact, model-agnostic validator. A union of eight Semgrep rules, derived from the three models' dominant regression signatures, catches 77.1% of AI-specific regressions while evaluating fewer than three percent of the available rule corpus. We also confirm that AI models exhibit statistically independent fixation and regression behavior, unlike human patchers, a finding that challenges standard patch review assumptions. Three forensic case studies ground the quantitative results in concrete code-level analysis and identify the shared mechanism underlying each model's failure pattern.

Details

Logo image