Why most AI pilots never reach production
The number everyone quotes is 95%. What it measures is not what most people think.
Pilots stall because nobody agreed what would count as working before the build started. Not because the model underperformed, not because the technology was immature, and almost never because of anything an engineer did. The research is unusually consistent on this, and it points at the same place every time: the decision made before any code existed.
Here is what the numbers actually say, and what they do not.
The 95% figure, and its caveat
MIT's Project NANDA published the number that now frames every conversation about this: 95% of generative AI pilots delivered no measurable profit-and-loss impact. The study reviewed over 300 publicly disclosed initiatives, ran 52 organisational interviews and surveyed 153 executives.
It gets quoted as 95% of AI projects fail. That is not what it measured. The finding is about P&L return, not about whether something shipped. A pilot can reach production, work exactly as specified, be used daily, and still register nothing on the earnings line — because it was scoped against a task rather than against a number anyone tracks.
That distinction matters, because the two failures have different fixes. A pilot that never shipped has a delivery problem. A pilot that shipped and returned nothing has a scoping problem, and no amount of better delivery will touch it.
The other figures fill in the picture:
- RAND (2024): more than 80% of AI projects fail to deliver intended business value — roughly twice the failure rate of conventional IT projects, which is not a flattering baseline to be double.
- IDC (2025): for every 33 proofs of concept an enterprise starts, about four reach production.
- S&P Global (2025): 42% of companies abandoned most of their AI initiatives, up from 17% the year before.
- BCG (2026): of 1,800 executives surveyed, 26% reported meaningful financial value from AI investment.
Read together, they describe a funnel that leaks at every stage, and leaks worst at the two ends: what got chosen at the start, and what got measured at the end.
Why pilots stall
The causes cited across MIT, RAND and BCG are almost identical, and RAND states the conclusion plainly: the failures are overwhelmingly organisational rather than technical.
Nobody wrote down what success would look like
This is the one that produces the 95%. A pilot is approved with a goal like improve customer response quality. It ships. Response quality is arguably improved. Nothing changes on any number the business tracks, and eighteen months later nobody can say whether it worked, so it is quietly not renewed.
The fix is not a better metric framework. It is one sentence written before the build: this number, currently at X, should be at Y by date Z, and if it is not, we stop. Teams resist writing that sentence, and the resistance is the signal — if you cannot name the number, the project is not ready to be funded.
The data was ready for a demo, not for production
Pilots run on a curated subset: cleaned, selected, owned by the team running the pilot. Production consumes real data governed by compliance rules and owned by several teams who did not agree to this.
Gartner's finding that 85% of AI projects fail on data quality is doing a lot of work in every article on this subject, and it is directionally right. But the practical version is simpler: if the pilot needed someone to prepare the data by hand, you have not tested the thing you are going to run.
It was never in anyone's workflow
A capability that requires leaving the tool where the work happens is not used, regardless of how good it is. This is the failure I have seen most often, and it is invisible in a demo — in a demo, someone opens the tool on purpose.
The cost is not the seconds spent switching. It is the context the person was holding, which does not survive the trip.
It was automation with a label on it
Some of these projects should never have been AI projects. In an assessment I ran across forty AI proposals inside a regulated European enterprise, roughly one in ten was doing something a rule, a database query or a lookup table would have done better — faster, cheaper, and without the governance overhead that attaches to a model.
Those are not failures of execution. They are correctly built solutions to a problem that was mislabelled at the start, and the label cost them a review process they never needed.
What the ones that ship do differently
The pattern in the surviving 5–26% is not exotic:
They scope against one number. Not a capability, not a theme. One measurable outcome with a baseline, a target and a date.
They set kill criteria before starting. A pilot with no defined way to fail does not end — it becomes permanent, consuming budget and producing nothing. The ability to name the condition under which you would stop is the clearest signal a project has been thought about properly.
They cut scope until it is uncomfortable. A narrow thing that works beats a broad thing people route around. That is not a compromise; on anything customer-facing it is usually the reason the numbers moved.
They cost the whole thing. A depressing share of AI business cases use sticker-price API costs with no allowance for evaluation, observability, governance, model updates, drift, security review or human quality assurance. When the real run cost appears, the return evaporates — and it was never there, it was an arithmetic error.
The question worth asking first
Before scoring, before feasibility, before anyone looks at the data:
Would a rule, a query or a lookup table do this job better? Not adequately — better. If the answer is yes, you have found automation wearing an AI label, and you have just saved yourself a governance process and a disappointment.
If the answer is no, the next question is the one most pilots never survive: what number changes, and by how much, and by when?
Neither question requires a data scientist. Both are cheaper to ask now than after eighteen months of funding.
Common questions
What percentage of AI pilots fail?
MIT's Project NANDA found 95% of generative AI pilots delivered no measurable profit-and-loss impact, and RAND put the broader AI project failure rate above 80%. The figures measure different things: NANDA measured return, not whether the pilot shipped.
Why do AI pilots fail to reach production?
Overwhelmingly for organisational rather than technical reasons: no agreed definition of success before the build, data that was ready for a demo rather than for production, no placement inside an existing workflow, and costings that ignored evaluation, governance and drift management.
Is the 95% figure reliable?
The 95% figure is well-sourced but routinely misquoted. MIT Project NANDA measured P&L return across 300 disclosed deployments, not the share of projects that shipped. A pilot can reach production, work as specified and still register nothing on the earnings line.
How do you stop an AI pilot from stalling?
Scope it against one number with a baseline, a target and a date, and agree the condition under which you would stop before you start. A pilot with no defined way to fail does not end; it becomes permanent.
Scoring 40 AI use cases in a regulated enterprise — the assessment this article describes, run across forty real proposals inside a regulated enterprise.