How to scope an AI pilot so it can actually ship
One number, one workflow, and a written condition for stopping.
A pilot that can ship has three things fixed before anyone writes code: one number it is trying to move, one workflow it lives inside, and a written condition under which you stop. Most pilots have none of them, which is why IDC found that for every 33 proofs of concept an enterprise starts, about four reach production.
One number, with a baseline
Not a capability. Not a theme. A number that already exists, that someone already looks at, currently sitting at a known value.
The sentence to write is: this number, currently X, should be Y by date Z. Three parts, all required. A target with no baseline cannot be evaluated; a baseline with no date never gets evaluated.
The resistance to writing that sentence is the useful part. If nobody will commit to the number, you have learned that the project is not ready to be funded — and you have learned it for the cost of a meeting rather than two quarters.
One workflow, named specifically
Which screen. Which person. What they were doing immediately before.
A capability that requires someone to leave the tool where the work happens does not get adopted, regardless of quality. The cost is not the seconds of switching — it is the context they were holding, which does not survive the trip.
If the answer to which screen is a new one we'll build, the pilot is now two projects and should be scoped as two.
Kill criteria, agreed in advance
A pilot with no defined way to fail does not end. It becomes permanent — consuming budget, occupying a slot on someone's roadmap, and generating nothing.
S&P Global found 42% of companies abandoned most of their AI initiatives in 2025, up from 17% the year before. Abandonment is what happens when nobody agreed what stopping would look like, so it arrives late and messily instead of on schedule.
Write it down: if the number has not moved by date Z, we stop and here is what we do with what we learned. Agreeing that while everyone is optimistic is considerably easier than agreeing it while everyone is defensive.
Then cut it until it is uncomfortable
The instinct is to cover every path, because a system that says I can't help with that feels like a failure. It is not.
On a change-management agent I scoped, the proposal covered the full process end to end. I cut it to the common path only, with everything unusual handed straight back to the existing form. That was the unpopular decision and the one the numbers came from: change creation went from fifteen minutes to seven, and policy compliance from 72% to 89%.
A narrow thing that works beats a broad thing people route around. The scope cut is not a compromise on the way to the real version — it is frequently the reason there is a real version.
What this costs you
Roughly a day of argument before anyone starts, and the loss of the comfortable feeling that comes from an ambitious scope.
What it buys is the ability to say, at the end, whether it worked — which 95% of generative AI pilots could not, according to MIT's assessment of 300 disclosed deployments.
Common questions
How do you scope an AI pilot?
Fix three things before any code: one number that already exists with a baseline and a date, one named workflow it lives inside, and a written condition under which you stop. Then cut the scope until it is uncomfortable.
What are kill criteria for an AI pilot?
The agreed condition under which you stop — typically that a named number has not moved by a named date. Agreeing it while everyone is optimistic is much easier than agreeing it while everyone is defensive, and a pilot without it becomes permanent rather than ending.
How small should an AI pilot be?
Small enough that it can be wrong cheaply. Covering the common path well and handing everything unusual back to the existing process usually outperforms attempting the full scope — an agent that is unreliable at the edges loses its credibility after the second bad answer.
Why do AI pilots stay stuck as proofs of concept?
Because nothing defined what finishing looked like. IDC found that for every 33 proofs of concept an enterprise starts, about four reach production. The rest are not cancelled; they simply never end.
Cutting enterprise change requests 15 to 7 minutes — the scope cut described here — 15 minutes to 7, and compliance from 72% to 89%.