Research

Twenty percent faster, nineteen percent slower

The developers in METR’s trial were not wrong about liking AI. They were wrong about what it did to the clock. The interesting number is not the slowdown, it is the thirty-nine point gap between what they measured and what they felt.

Published
25 August 2026
Reading time
4 minutes
Primary source
METR, July 2025
Design
Randomised trial
Bottom line
Trust the stopwatch
The trial

Sixteen experienced developers, in their own repositories.

Most claims about AI coding productivity come from benchmarks, vendor telemetry or self-report. METR did something harder. They took experienced open-source developers, gave them real issues from repositories they already maintained, and randomised which tasks they were allowed to use AI assistance on.

That design matters. These were not people learning a codebase, where an assistant is obviously useful. They were people who already held the context in their heads, working on code they had written, which is the exact situation most senior engineers are in most of the time.

The result was that tasks took longer with AI assistance available than without it.

39 points

between perception and measurement. Developers finished 19% slower, and afterwards estimated that AI had made them 20% faster.

METR · randomised trial, July 2025
The gap

The slowdown is interesting. The misperception is the finding.

It would be easy to read this study as "AI makes developers slower" and stop. That reading is fragile: the study was small, the tooling in early 2025 was not the tooling of today, and other contexts plainly do show gains.

What does not depend on any of that is the gap. Everyone in the trial had just done the work. They still could not tell which direction the effect ran, by a margin larger than the effect itself.

That has an uncomfortable consequence for how most organisations decide about tooling. If the people doing the work cannot feel a 19% difference, then a survey asking them whether the tool helped is not measuring the tool. It is measuring how the tool felt.

What is happening

Waiting and reading do not feel like work.

The plausible mechanism is about where the time goes rather than how much there is. With an assistant, a task splits into prompting, waiting, reading, deciding, and correcting. Each piece is short. None of it feels like effort in the way that staring at a problem does.

  • Writing a prompt feels like progress, because something appears.
  • Waiting feels like nothing, because you are not doing anything.
  • Reading generated code feels fast, because it is well-formatted and confident.
  • Correcting it feels like a small detour rather than a cost.

Add those up across a day and the stopwatch says one thing while the memory says another. The memory is recording effort, and effort genuinely went down.

What it does not mean

Stop using the tools is the wrong conclusion.

We use them. So does every engineer on our bench. The conclusion we draw is narrower and more useful: do not let anybody’s impression of a tool substitute for a measurement of the thing you actually care about.

Measure cycle time, not sentiment

Time from problem stated to change in production, per change, tracked before and after. It is unglamorous and it is the only number that cannot be argued with.

Expect the gain where the context is missing

The trial found a slowdown among people who already knew the code. The inverse is worth testing in your own shop: assistance tends to pay where somebody is new to an area, and to cost where they are not.

Watch the handoffs, not the keystrokes

In most delivery timelines we are asked to look at, the waiting between people dwarfs anything happening at the keyboard. Doubling typing speed in a process where 60% of the calendar is queueing changes almost nothing, and no tool fixes a queue.

Sources

Where every number here came from.

Cited in this piece

  1. METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. metr.org
  2. METR research index, for the study’s methodology and subsequent commentary. metr.org
Next step

Measure it before you scale it.

Thirty minutes on where your cycle time is really going, and whether more tooling or fewer handoffs is the cheaper fix.