Veracode
Confidence 0.70 · 1 source · last confirmed 2026-08-30
An application security testing vendor, present in this wiki as the publisher of the 2025 GenAI Code Security Report — the largest systematic scan of AI-generated code security in the corpus.
The study asked over 100 large language models, spanning several years of sizes and vintages, to complete 80 curated coding tasks in Java, Python, C# and JavaScript, each designed so that a secure and an insecure completion were both plausible. 45% of samples introduced OWASP Top 10 vulnerabilities, with a near two-fold spread by language (Java 72%, Python 38%) and XSS failing in 86% of relevant samples.
Its most consequential finding is the negative one: security performance was “flat, regardless of model size or training sophistication.” Newer and larger models were “no better.” This is the corpus’s primary evidence that security is not on the capability curve — the one dimension that has not improved as SWE-bench resolution ran from 1.96% to 87%.
Veracode sells the remediation for the problem it measures, which is a real interest to hold in view. The flat-scaling result is corroborated in shape by independent academic work (Spracklen et al.), which is why it is treated as credible here. The announcement was authored by Jens Wessling, Veracode’s CTO.