Engineering

A year of model progress did nothing for code security

Veracode ran the same benchmark twice, a year apart. The models got dramatically better at writing code that works. The share of that code carrying a real security flaw barely moved. That gap is not a model problem you can wait out.

Published
11 August 2026
Reading time
5 minutes
Primary source
Veracode, July 2026
Sample
100+ models
Bottom line
Review is the bottleneck
The finding

Fifty-five percent, then fifty-six.

Veracode runs a benchmark that asks language models to complete realistic coding tasks where an insecure answer is available and tempting, then checks whether the code they produce carries a known vulnerability class. They ran it in 2025. They ran it again in 2026 across more than a hundred models.

Between those two runs, the models changed enormously. Context windows grew, reasoning got better, agentic coding went from a demo to a default. The security pass rate went from 55% to 56%.

44%

of AI code generation tasks introduced a risky security vulnerability. The best model in the field managed 68%, which still means roughly a third of its output needed fixing.

Veracode · 2026 GenAI Code Security Report

The useful way to read that is not as an indictment of the tools. It is a statement about what the tools were optimised for. A year of training pressure went into producing code that runs, compiles, passes the test and satisfies the person reading it. None of those signals punish a vulnerability, because a vulnerable function behaves exactly like a safe one until somebody attacks it.

Why it stalled

Nothing in the loop was ever asking for security.

Consider what a model sees when it learns to code. It sees public repositories, which contain a great deal of code that works and a great deal of code that is quietly unsafe. It sees a reward for output that a developer accepts. It sees test suites, which check behaviour rather than exposure.

A vulnerability is defined by what happens under an input nobody wrote a test for. It is invisible to every signal in that loop, which is exactly why a year of progress on every other axis left it where it was.

There is a second effect underneath. In METR’s randomised trial, experienced developers working in their own repositories were 19% slower with AI assistance while believing they had been 20% faster. If your sense of how carefully you reviewed something is that badly calibrated, the review step is weaker than it feels, at exactly the moment it is carrying more load.

Everything before the review step is fast and confident. The review step is the only thing between a model and your production environment.
It is not uniform

Some vulnerability classes are nearly solved. Others are not close.

The average hides a wide spread. In the 2026 report, models handled SQL injection at an 83% pass rate and cryptographic algorithm choice at 87%. Both are well represented in training data, well documented, and mostly a matter of using the parameterised call instead of the string concatenation.

Cross-site scripting came in at 15%. Log injection at 12%.

That pattern is not random. The classes the models handle well are the ones with a single correct local fix. The classes they fail are the ones where safety depends on context the model cannot see from the function it was asked to write: where this string will be rendered, who controls it, whether it has already been escaped once upstream.

Which is another way of saying the failures cluster exactly where knowing the system matters more than knowing the language.

What to do about it

Three things, none of them about the tool.

Put a name against the merge, not the file

Most teams already know who wrote a module. Fewer can say who is accountable for what it does in production six months later. When the writing is cheap, ownership is the only thing left that scales, and it has to belong to a person rather than a rota.

Make the review adversarial about context, not style

Reviewing AI output for readability is wasted effort; it is already readable, which is part of the problem. The useful review questions are the ones the model could not have answered: where does this input come from, what happens if it is hostile, and has this value already been trusted somewhere upstream.

Scan for the classes the models are bad at

If injection through rendering and logging is where the pass rate collapses, that is where the static analysis budget belongs. Tuning your pipeline to the known weak classes costs a day and catches the errors the generation step is statistically likely to make.

None of this requires abandoning AI assistance, and we would not suggest it. Our own engineers use these tools daily. The point is narrower: the writing got cheap and the reviewing did not, so the reviewing is now the entire job.

Sources

Where every number here came from.

Cited in this piece

  1. Veracode, 2026 GenAI Code Security Report, published 28 July 2026. veracode.com
  2. Veracode, 2025 GenAI Code Security Report, the prior year’s run of the same benchmark. veracode.com
  3. METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, July 2025. metr.org
Next step

Have somebody senior read what the model wrote.

Thirty minutes to talk about where AI-generated code is entering your pipeline and who is currently accountable for it.