Currently free during beta - premium features coming soon. Subscribe now to lock in early access.

arXiv: From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts

AI_SAFETY AI Security & Safety · · arxiv_cscr

AI Analysis

This paper, published on arXiv, presents an independent reproducibility study of large language model (LLM) and agent-driven tools designed to automatically validate software vulnerabilities. The study found that while these AI systems can generate plausible security artifacts, their outputs are often not directly runnable or verifiable in real-world environments, with a significant portion failing to reproduce the claimed vulnerability validation results. The core change is a documented evidence base showing that current AI-driven security tooling lacks the reliability and traceability required for formal compliance evidence.

The primary audience is organizations in regulated sectors that rely on automated security testing, including financial services, healthcare, critical infrastructure, and any software vendor subject to EU cybersecurity frameworks like the Cyber Resilience Act (CRA) or NIS2. Compliance teams and DevSecOps leaders who are considering or already using LLM-based vulnerability scanners must treat their outputs as unverified claims, not as auditable findings. This affects risk assessments, audit trails, and any evidence submitted to regulators or certification bodies.

Compliance teams should immediately update their vendor due diligence and internal validation procedures. Specifically, they must require human-in-the-loop verification for any AI-generated vulnerability report, mandate that all artifacts be re-run in a controlled sandbox, and document the reproducibility rate of their chosen tools. Additionally, they should add a clause in procurement contracts holding AI vendors accountable for the verifiability of their outputs, and prepare to explain to auditors why AI-generated evidence was not used as standalone proof of security compliance.

Get notified about AI_SAFETY changes

Subscribe to our free weekly digest covering 24 compliance frameworks.