Why most enterprise AI pilots produce nothing
The blocker was never model quality. In the study that produced the widely quoted 95% figure, the projects that worked differed from the ones that did not in a single respect, and it had nothing to do with which model they picked.
Ninety-five percent, and the part everyone skips.
In 2025 a group at MIT looked at roughly three hundred public enterprise generative AI deployments and asked a deliberately boring question: did this show up in the profit and loss statement. For about 95% of them, the answer was no.
That figure travelled a long way, usually as a headline about AI being overhyped. The more useful part of the report is the comparison inside it, which almost nobody quotes.
Initiatives built with an external specialist partner succeeded roughly twice as often as those built entirely in-house. Same models, same market, same year.
Gartner, looking forward rather than back, expects more than 40% of agentic AI projects to be cancelled before the end of 2027, citing unclear business value and escalating costs rather than technical limits.
Two different methods, the same conclusion: the constraint is organisational.
Proximity to the workflow, and somebody accountable for it.
The projects that produced a result shared a shape. Somebody with engineering authority sat close enough to the actual work to see where it broke, and had the standing to change the process rather than only the software.
The ones that did not produce a result were usually competent. The model was fine. The demo was impressive. What was missing was the unglamorous middle: the part where you find out that the operations team has three exceptions to the rule nobody documented, that the data arrives in two formats, and that the step the pilot automated was not the step that cost anybody time.
That work is not technically difficult. It is just nobody’s job in the standard structure, where the vendor builds to a specification and the internal team owns a roadmap.
A pilot is designed to produce a decision, not a system.
Run a pilot and you get a demo, a deck, and a meeting. All three are real outputs. None of them is a thing that runs on a Tuesday when the person who built it is on leave.
The failure is structural rather than lazy. A pilot is scoped to be cheap, which means scoped to avoid the integration, the permissions model, the exception handling and the support arrangement, which are precisely the things that determine whether it survives. So the pilot succeeds on its own terms and dies on contact with production, and the organisation concludes that the technology was not ready.
Deploy the engineer, not the proof of concept.
Start in production, small
The first thing we ship goes into a real environment with real users, even if it does almost nothing. A narrow thing that runs teaches you more in a week than a broad thing that demos teaches you in a quarter.
Put the engineer where the problem is
Not in a delivery pod with a liaison. In your standup, reading your support tickets, asking your operations lead why a workflow exists. That is where requirements actually live, and it is the specific step the 95% skipped.
Name the metric before anything is built
One number, agreed in writing, that this work is supposed to move. If nobody can name it on the first call, the honest answer is that the project is not ready, and we would rather say that than bill for a quarter finding out.
Keep the person who built it
Most of what kills a working system is the handover. The engineer who built it is the engineer you call when it breaks, which is less a service promise than an admission that written handovers do not carry the reasoning.
The 95% figure deserves more scepticism than it usually gets.
It came from a working paper rather than a peer-reviewed study, the sample was drawn from public deployments rather than a random selection, and "no measurable P&L impact" is a high bar for any initiative inside its first year. Several credible people have said so.
We quote it anyway because the internal comparison survives those objections. Whatever the true failure rate is, the in-house and partnered projects in that sample were measured the same way, against the same bar, in the same year. The ratio between them is the part that carries information, and it happens to agree with what Gartner expects and with what we see in the accounts that come to us after a stalled pilot.
Where every number here came from.
Cited in this piece
- MIT NANDA, The GenAI Divide: State of AI in Business, August 2025, covering roughly 300 public enterprise deployments. Widely reported; see for example Fortune’s coverage.
- Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, press release, 25 June 2025. gartner.com
- Veracode, 2026 GenAI Code Security Report, 28 July 2026. veracode.com
Bring us the pilot that stalled.
Thirty minutes on why it stopped, and a straight answer about whether it is worth restarting or worth killing.