AI Reality Check

An honest read on which of your AI ideas will actually ship.

You bring the roadmap. I score every idea against the same six questions and give you a build order you can defend.

One week  /  fixed scope  /  from €1,500

Who this is for

You have five to twenty AI ideas and no way to compare them.

They came from a board meeting, a competitor's website, a vendor pitch and an offsite. Everyone has an opinion about which ones matter and none of the opinions agree.

You don't need a strategy deck. You need someone with no stake in the answer to tell you which of these are real.

What you get

Six deliverables, all yours whatever happens next.

01

A scored table

Every idea marked real, theatre or flagged, with the reasoning written beside it so anyone can audit the call.

02

Effort and impact

Where each surviving idea sits, with the number it should move and where that number stands today — so the sequencing argument is visible rather than asserted. Where the number doesn't exist yet, that's the first finding.

03

A build order

The three to five worth doing, in the order I'd do them, and what to do first.

04

A kill list

What to stop, and what to do instead — often a rules engine, an off-the-shelf tool, or nothing.

05

Risk and governance

Where each use case lands under the EU AI Act, and what that means before you build.

06

A live readout

I present all of it to the people who have to act on it, and answer for every score.

Sample output — illustrative

Auto-summarise support ticketsREAL
Rules engine, relabelled as AITHEATRE
Search over your own documentsREAL
Chatbot reading a FAQTHEATRE
Churn prediction modelFLAGGED — NO DATA

Flagged means the idea is real but blocked until something changes — usually data that doesn't exist yet, or an owner nobody has named.

Sample output — effort against impact

LOW EFFORT HIGH EFFORT LOW IMPACT HIGH IMPACT DO FIRST DO NEXT MAYBE DON'T Ticket summarisation Doc search Draft generation FAQ chatbot Relabelled rules engine Churn model — flagged

The sequencing argument becomes visible rather than asserted. Filled dots are real, dashed are theatre, outlined are flagged.

How the scoring works

The same six questions, asked in the same order, every time.

Two gates first. Anything legally prohibited stops there. Anything with no realistic path to the data gets parked rather than scored. Then every surviving idea gets the same treatment, so the answer doesn't depend on who pitched it or how senior they are.

First — two gates

Fail either and it never reaches a score.

Is it legal?Prohibited under the EU AI Act stops here
Is the data reachable?No path to data means parked, not scored

Question 1 — scored 0–3

Does this need AI at all?

Would a rule, a query or a lookup table do this better? If so, that's a recommendation worth paying for. Boring technology that works beats a model that impresses.

Question 2 — scored 0–3

Does the data exist?WEIGHTED ×1.5

Not “could we collect it” — does it exist today, at volume, clean enough, with the right to use it. For generative use cases the question changes: is there a test set and governed context, rather than training labels.

Question 3 — scored 0–3

What's the actual upside?WEIGHTED ×1.5

In euros or hours, against a documented baseline, as a range rather than a single confident number. If nobody can size it, that tells you something.

Question 4 — scored 0–3

What breaks if it's wrong?

Wrong answers are guaranteed. If a wrong answer moves money, breaks a regulation or harms a customer, the bar moves and the governance work comes first.

Question 5 — scored 0–3

Build or buy?

Plenty of genuine AI is a commodity now. Building it yourself is a choice that needs defending — and it changes your legal position, since building can make you a provider under the EU AI Act where buying leaves you a deployer.

Question 6 — scored 0–3

Who owns it after launch?

Models drift and quietly get worse. If nobody owns the monitoring and the retraining, you're not shipping a feature, you're shipping a liability.

Finally — one verdict

Scores resolve to one of three.

REALWorth building. Enters the build order.
THEATREKill it. Usually a rule, a tool, or nothing.
FLAGGEDReal, but blocked until something changes.

The rubric is the easy part.

Six questions you could ask yourselves. Most teams don't, and when they do the answers come back wrong — not because the questions are hard, but because of who's answering them.

The person who proposed the use case is scoring their own use case. The executive who announced the AI programme is sitting in the room where you decide it's theatre. The data team says the data exists, because saying otherwise means admitting a migration didn't land. None of these people are lying. They just can't give you a clean answer about their own work.

I have no stake in which way it goes. That's most of what you're paying for.

Where the money goesMost AI budget lands in sales and marketing, where the measured return is lowest. Scores tell you what's real. They don't tell you what to fund first.
Surviving contact with the workMost pilots die between demo and production. Judging which will hold up inside a real workflow is a different call from scoring the idea.
What it costs to say itSome theatre is expensive to name out loud, especially when someone senior announced it. Knowing when to say it anyway, and how, is judgement.

The rubric is calibrated against my own engagements and against the NIST AI risk framework and the EU AI Act. I'll tell you which parts are established practice and which are my judgment — you should know the difference.

What it costs

Fixed price, so the scope can't drift on you.

Startup
€1,500

Up to ten use cases

  • Intake questionnaire
  • One working call, about ninety minutes
  • Written report and build order
  • Live readout
Enterprise & mid-market
€3,500

Up to twenty use cases

  • Everything in the startup tier
  • Stakeholder interviews
  • Deeper governance and EU AI Act section
  • Readout built for leadership
“I've worked with a lot of product people in AI and R&D. Fanni is the one who asks whether the thing should be built at all. She'd rather kill a bad idea in week one than let it eat a quarter.”
Csaba Molnar  /  AI Robotics
“Working with Fanni was straightforward in a way large-organisation projects usually aren't. She asked the awkward questions early, which meant we didn't discover them late.”
Akos Hulej  /  ExxonMobil

Where it comes from

Built where saying no had consequences.

40AI proposals scored
1 in 10were not actually AI

I built this rubric inside a regulated European enterprise and ran forty AI proposals through it. Every score had to survive security review, legal and a works council.

Saying no there had consequences. That is where the six questions come from.

Most AI assessment frameworks are written by people selling AI. This one was built by someone who had to say no inside an organisation that had already announced it was saying yes.

Read that case study

Questions

Before you book.

Because it takes real work and the answer has value on its own. The kill list alone usually saves more than the fee. If you want a free directional read first, the scorecard takes four minutes.

Then I've saved you a quarter. That's a legitimate result, and it happens.

If you sell into the EU it applies to you. If you don't, I'll cut that section and go deeper elsewhere.

An intake questionnaire, one call of about ninety minutes, and an hour for the readout.

Yes, as standard.

Then we scope it as two rounds, or narrow to the ones with budget attached. I'll tell you which on the intro call.

Your turn

Bring me the list nobody's willing to question.

Twenty minutes, no pitch. If I'm not the right person, I'll say so and point you somewhere better.

Try the free scorecard first