Case study

Forty AI use cases. Roughly one in ten wasn't AI.

The assessment framework behind the Reality Check, built inside an organisation where getting it wrong has consequences.

AI governanceRegulated enterpriseClient not named
40use cases assessed and classified
1 in 10filtered out as not genuinely AI
509 hrsa month identified across the quantified cases

The situation

A large regulated European enterprise had AI proposals arriving from every direction — business units, vendors, and internal teams who had all been told to have an AI plan.

Nobody could compare them. Each proposal was written in its own language, sized by its own sponsor, and judged mostly on who was asking. Some were substantial. Some were existing systems with a new label. The organisation had no consistent way to tell the difference, and in a regulated environment the cost of getting that wrong isn't just wasted budget.

I can't name the client. What I can describe is the method, because the method is the transferable part.

What happened

I built an assessment framework and ran forty use cases through it.

Every proposal got the same questions in the same order, scored the same way, with the reasoning written down beside each score. That last part matters more than it sounds: a score you can audit survives a conversation with a sceptical stakeholder, and a score you can't is just an opinion with a number attached.

The framework had to hold up in rooms that don't accept hand-waving — security review, legal, and a works council. It also had to sit inside a live regulatory boundary, which meant classification wasn't optional and “we'll deal with compliance later” was never an available answer.

I fed the product and technical assessment into that process. I didn't do the legal work, and I'm careful to say so.

The result

Forty use cases assessed and classified. Roughly one in ten filtered out as not genuinely AI — proposals a rules engine, an existing system or a query already handled. Around 509 hours a month of identified saving across the two dozen use cases that were quantified.

The rejections were the valuable part. Each one was a project that didn't consume a quarter of engineering time to arrive somewhere the organisation already was.

What this means if you're reading it

This is where the Reality Check comes from. The rubric I use with clients is the same instrument, adapted from an environment where every score had to be defensible to someone whose job was to challenge it.

Most AI assessment frameworks are written by people selling AI. This one was built by someone who had to say no inside an organisation that had already announced it was saying yes.

What was hard

Scoring somebody's proposal low is scoring somebody's budget, headcount and internal case low. In an organisation where every unit had been told to have an AI plan, a framework that filters proposals out is a framework that makes enemies unless the reasoning survives contact with the person it disadvantages.

That is why the written reasoning mattered more than the score. A number on its own invites an argument about the number. A number with the reasoning beside it moves the conversation to the reasoning, which is where it should have been.

The other difficulty was the ones that were nearly AI. Filtering out an obvious rebrand is easy. Deciding about a proposal that would genuinely benefit from a model but would benefit more from fixing the process underneath it is the judgment the framework had to support rather than replace.

Common questions

How do you tell an AI project from a project with AI written on it?

Ask whether a rule, a database query or a lookup table would do the job better — not merely adequately. Across forty proposals, roughly one in ten failed that test: existing systems with a new label.

What makes an assessment framework hold up?

The reasoning written down beside each score. A score you can audit survives a conversation with a sceptical stakeholder; a score you cannot is an opinion with a number attached.

Why score every proposal the same way?

Because otherwise they are judged on who is asking. Each proposal arrives written in its own language and sized by its own sponsor, which makes comparison impossible until the same questions are asked in the same order.

Your turn

Bring me the list nobody's willing to question.

Twenty minutes, no pitch. If I'm not the right person, I'll say so and point you somewhere better.

See the other projects