Every AI initiative I’ve been asked to rescue had the same thing missing, and it was never the model, the engineers or the executive support. It was a number.
Is it true that 95% of AI pilots fail?
Sizing an AI feature means putting a number on it before any code exists, by answering three questions: what this is worth if it works, what happens if we do nothing, and what the baseline is today. Most teams skip all three. In BCG’s 2026 survey, only 14% of companies had clearly defined P&L impact for all their AI initiatives.
Nobody had written down what the thing was worth if it worked — so there was nothing to measure against. So there was no way to tell whether it was working, no way to decide when to stop, and no way to defend the budget when someone senior got nervous. The project drifted until it was quietly absorbed into something else.
This is the most boring failure mode in AI product work and it’s by far the most common.
The evidence, and one number you should stop quoting
Let me start with the statistic everyone uses, because I think using it carelessly is part of the problem.
You have seen the claim that 95% of enterprise AI pilots fail. It comes from The GenAI Divide: State of AI in Business 2025, a July 2025 report from Project NANDA, branded MIT NANDA.
What it actually found is narrower than the headline. Roughly 95% of organisations in its evidence base hadn’t achieved measurable P&L impact from generative AI initiatives within the study’s criteria and time window. It didn’t find that 95% of AI projects are technically broken or permanently unsuccessful. The study defined success as deployment beyond pilot with measurable KPIs, and assessed ROI six months after the pilot, which will undercount longer enterprise implementations.
And how was the evidence base built? Structured interviews with representatives from 52 organisations, survey responses from 153 senior leaders recruited at four industry conferences, and analysis of 300-plus publicly disclosed initiatives. Note the self-selection in a conference sample.
There’s a citation problem on top of that. The PDF states 52 interviews and 153 leaders. Fortune, which put the story into the mainstream in August 2025, reported 150 interviews and a survey of 350 employees. Those are different studies as described, and the second version is the one most people are repeating.
And the report was produced by a project that builds agentic AI infrastructure, and concluded that agentic infrastructure is the path across the divide it identified. That isn’t evidence of bad faith. It’s the kind of thing that should make you read more slowly.
I’m spending this long on it for a reason that matters to the rest of the article. The 95% number is used, constantly, as an argument for buying something. If you walk into a budget conversation quoting it, you’re borrowing authority you haven’t checked, and the first person who has read the PDF will take the room off you.
What the more careful sources say
Strip the headline away and the underlying picture still holds, which is the interesting part.
McKinsey’s 2025 State of AI surveyed 1,993 participants. 88% said their organisations regularly used AI in at least one function, while only 39% reported any enterprise-wide EBIT impact. Their phase data shows 32% experimenting, 31% piloting, 30% scaling and 7% fully scaled.
The UK’s Office for National Statistics, analysing in July 2026, found 35% of businesses with at least 10 employees used at least one AI technology, and among adopters only 10% reported extensive use. That’s a national statistics office rather than a vendor, which makes it the source I’d take into a boardroom.
BCG’s July 2026 survey of 152 chief executives at companies with revenues of at least $500 million found more than half naming the link between AI initiatives and the P&L as a key barrier, with only 14% having clearly defined P&L impact for all their AI initiatives.
Four different methodologies, one consistent finding: broad use, shallow financial impact, and a widespread inability to connect the two. None of them requires the 95% headline to make the point.
So the sizing question is the whole job
If more than half of large-company chief executives say linking AI to the P&L is a barrier, then sizing is not a preliminary step before the real work. It’s the work most people are failing at.
Here’s what I do.
Find the baseline before you promise anything. What does this cost you today, in time, money, error rate, or whatever your unit is? If nobody knows, that’s your first deliverable and it’s usually a week. You’ll want to skip it, because it’s unglamorous. Then you launch, and you can’t prove an improvement because you never measured the before. A pilot without a baseline produces anecdotes.
Write the arithmetic on one page. Volume times unit cost times expected improvement. Rough is fine. You’re not after precision. You’re making your assumptions visible so somebody can disagree with a specific one. A number you can argue with beats a conviction you can’t.
Name the do-nothing case honestly. What happens if we skip this? Sometimes the answer is a competitor eats us. More often the answer is nothing much, and that’s useful to know in month one rather than month nine.
Decide what would make you stop. Agree the number before you start, because afterwards you’re invested and so is everyone else. It’s the hardest thing to get a room to commit to. It’s also what stops a two-year zombie project.
Then check whether the value needs AI at all. A meaningful share of what gets scoped as an AI feature is a reporting problem, a process problem, or three lines of SQL. Finding that out costs a day. I’ve never regretted asking.
Where the arithmetic goes wrong
Two mistakes, and I see both constantly.
The first is sizing the ceiling rather than the realistic case. If your model is right 80% of the time and a human still checks the other 20%, you don’t get 80% of the benefit. You get what’s left after review time, and review time is often the dominant cost. Size the workflow, not the model. This is the same problem I wrote about in coverage isn’t value, arriving from a different direction.
The second is leaving out the cost of being wrong. Every error has a downstream price: a rework loop, a support ticket, a bad decision, a regulatory exposure. If your arithmetic only counts wins, you’ve written a sales deck.
What sizing gets you that nothing else does
A defensible stop condition. A number to report against that somebody senior already cares about. And the ability to say no to the next three things that arrive without one, which over a year beats any individual feature you’ll ship.
It also changes the build-versus-buy conversation completely, because you can’t compare options against an undefined benefit. I wrote about that separately in build or buy, and sizing is the step that has to happen first.
Common questions
Is it true that 95% of AI pilots fail?
Not as usually stated. The figure comes from The GenAI Divide, a July 2025 Project NANDA report, which found roughly 95% of organisations in its evidence base hadn't achieved measurable P&L impact within the study's criteria and time window. It didn't find that 95% of AI projects are broken or permanently unsuccessful. Success was defined as deployment beyond pilot with measurable KPIs, and ROI was assessed six months after the pilot, which undercounts longer enterprise implementations.
What is the evidence base behind that report?
Structured interviews with representatives from 52 organisations, survey responses from 153 senior leaders recruited at four industry conferences, and analysis of more than 300 publicly disclosed initiatives. The conference recruitment makes the sample self-selected. Fortune's August 2025 coverage described 150 interviews and a survey of 350 employees, which doesn't match the report's own figures.
What do more reliable sources say about AI value?
They agree on the shape without the headline. McKinsey's 2025 State of AI found 88% of 1,993 respondents used AI in at least one function while only 39% reported enterprise-wide EBIT impact. The UK Office for National Statistics found in July 2026 that 35% of businesses with at least 10 employees used AI, and only 10% of adopters reported extensive use. BCG's 2026 survey of 152 chief executives found only 14% had clearly defined P&L impact for all their AI initiatives.
How do you size an AI feature before building it?
Establish the current baseline, write the arithmetic on one page so the assumptions are visible and arguable, name what happens if you do nothing, agree in advance what result would make you stop, and check whether the value actually requires AI.
Why do AI pilots fail to show financial impact?
Most often because no baseline was recorded before the pilot started, so improvement can't be demonstrated afterwards. Sizing the model's accuracy rather than the end-to-end workflow, and omitting the cost of wrong answers, both inflate the expected value in ways that show up later as a missing return.
What is a baseline and why does it matter?
The measurement of how the work performs today, before any AI is introduced — cost, time, error rate, or whatever unit the value is claimed in. Without it a pilot can produce anecdotes but not evidence, and any improvement claim afterwards is unfalsifiable.
How do you calculate ROI on an AI project?
Volume times unit cost times expected improvement, against the baseline you recorded before starting, minus the cost of handling wrong answers. Precision matters less than making each assumption visible so somebody can dispute a specific one. Size the end-to-end workflow rather than the model's accuracy, since review time on the cases the model gets wrong is often the dominant cost.
When should you stop an AI project?
Agree the stop condition before you start, while nobody is invested yet. A number set in advance is the most reliable protection against a project that continues for two years because cancelling it would embarrass someone.
Does every AI use case actually need AI?
No, and checking costs about a day. A meaningful share of what gets scoped as an AI feature turns out to be a reporting problem, a process problem, or a straightforward database query.
I’m an AI product manager working across fintech, SaaS, and regulated enterprise — currently leading AI and workflow product at T-Systems International. If you’re building AI governance into a product right now and want to compare notes, I’m at csincsakf@gmail.com or on LinkedIn.