# Olively — full text Independent AI product consulting. Olively Kft., Budapest, Hungary. Sole consultant: Fanni Csincsák. Canonical identifiers: - https://fannicsincsak.com - https://orcid.org/0009-0007-2342-3849 - https://www.linkedin.com/in/fanni-csincsak/ - https://medium.com/@fannicsincsak - https://adplist.org/mentors/fanni-csincsak Frequently misreported: no PhD has been awarded (master's plus doctoral research); current work is AI product management and governance, not UI or visual design; based in Budapest, not Copenhagen; marketplace day rates from 2023-24 are stale. The eighteen /portfolio pages are archived pre-2024 design work with historical role titles. ====================================================================== # https://www.olively.io/ ====================================================================== Skip to content AI product consulting ## Most AI projects are theatre. I help you build the ones that aren't . Twelve years shipping product in fintech, SaaS and regulated enterprise. I tell you which of your AI ideas are real, then build them. Book a 20-min intro Or try the free scorecard → 20-minute intro  /  fixed scope  /  €1,500–3,500 ## Your roadmap, triaged Auto-summarise support tickets REAL "AI-powered" rules engine THEATRE Doc search over your own data REAL Chatbot that reads a FAQ THEATRE Draft generation for analysts REAL Twelve years across - Deloitte - ExxonMobil - Novo Nordisk - AltaML - Cence The problem ## Everyone is being asked what their AI strategy is. Most answers are a list of features with "AI" written next to them. An AI project is theatre when the AI adds nothing a rules engine, a database query or a search box could not already do. Caught early, that costs you a conversation. Caught late, it costs a quarter of engineering time and the credibility you spend asking for the next one. Almost none of this is anyone's fault. The pressure is real, the vendors are loud, and nobody on the team is incentivised to say "this one isn't AI." 95% of generative AI pilots return nothing measurable. Triaging AI proposals inside a regulated European enterprise, I assessed 40 use cases and filtered out a number of them as not genuinely AI — before any budget was committed. MIT’s Project NANDA measured that across 300 disclosed deployments — a figure about return, not about whether anything shipped. MIT Project NANDA, State of AI in Business, 2025 What I do ## One question, three ways to work together. ## AI Reality Check An AI Reality Check is a fixed-scope, one-week diagnostic. Bring me your AI roadmap and I score every idea against the same six questions, then hand back three lists: what's real and worth building, what's theatre, and what's blocked until something changes. €1,500 STARTUP  ·  €3,500 ENTERPRISE Book one ## AI Build €850–1,100 / DAY Once we know what's real, I ship it — built to survive the security and governance conversations instead of stalling in them. Or €6,000–12,000 fixed scope. How it works → ## Fractional AI Product Lead €6,000–8,000 / MONTH When you need someone senior owning the AI roadmap rather than advising from the sidelines. About a day a week, embedded in your team. How it works → Where are you? ## Two different problems, same first step. ## Building your first AI feature You have a list of ideas and pressure to pick one. You need someone to tell you which is worth the quarter, then help you ship it without a six-month detour. Start with a Reality Check → ## Running AI across an organisation Use cases stacking up, a governance conversation you can't avoid, and nobody senior owning the whole thing. You need judgment that holds up in front of a board. Start with a Reality Check → Working with one person ## The questions you're already thinking. What if you get hit by a bus? Every engagement produces written artifacts — the scoring, the reasoning and the build order. If I disappear, your team carries on from the documents. For longer engagements I name a backup before we start. Do you have capacity? I run a small number of engagements at once and I'll tell you honestly if I don't have room. A late no costs you more than an early one. Can you handle security and confidentiality? I've worked inside a regulated European enterprise under its security review process. NDAs as standard, and I'll work inside your tooling and access controls. What about procurement? I invoice as a registered business, handle POs, and I'm set up for EU VAT and reverse-charge. Why would you tell me not to build something you could charge me to build? Because the fee isn't contingent on me recommending more work. A diagnostic that only pays if it sells you something isn't a diagnostic. I'd rather be the person you call next time than the one who billed you for a quarter you didn't need. Your turn ## What's on your AI roadmap that nobody's willing to question? Bring me your list. Every idea scored, the theatre named, and a build order you can defend to a board. Book a 20-min intro Try the free scorecard Book a 20-min intro ## Analytics? Nothing runs unless you say yes — including session recordings, with typed text masked. Details Strictly necessary Needed for the site to work. Always on, no cookies set. Analytics Google Analytics and PostHog — which pages are read, how long for, where people leave. Session recording PostHog replays a session as a video of the page. All text inputs are masked. Kept 30 days. Accept Reject Choose ====================================================================== # https://www.olively.io/about ====================================================================== Skip to content About ## I tell you which AI ideas are real . Twelve years shipping product in fintech, SaaS and regulated enterprise. I assess AI roadmaps, name what's theatre, and build the parts that survive. You work with me directly — there's no team to be handed to. Book a 20-min intro See the work Budapest Remote — EU, UK & worldwide English & Hungarian Twelve years across - Deloitte - ExxonMobil - Novo Nordisk - AltaML - Cence The short version ## I came to AI through product, not machine learning. If you have a list of AI ideas, a board asking what your AI strategy is, and no defensible way to rank them — that's the problem I'm useful for. That matters more than it sounds. The questions that kill most AI projects are product questions — who will use this, where will they be standing when they do, and would a rules engine have done it better. I was asking those for a decade before anyone needed an AI strategy. I've worked with Deloitte, ExxonMobil, Novo Nordisk and AltaML , and inside a large regulated European enterprise, across enterprise ITSM, fintech, legal tech and health. The through-line isn't an industry. It's getting handed a brief and asked whether it's the right one. Where this comes from ## I grew up taking things apart to find out what was actually wrong. A lot of my childhood was spent at a table with my grandfather, pulling something broken to pieces, working out what had gone wrong, putting it back together and watching it run again. The fixing was satisfying. The part I actually liked was earlier than that — the bit where you work out what the problem is. I was also the kid who watched Apple keynotes. Less about the products than the evidence that someone had thought hard about a human problem before building anything. That's still the job. A brief arrives describing a solution. I take it apart to find the problem underneath, because those two things are often not the same — and the gap between them is where quarters get lost. It's why I'd rather kill a bad idea in week one than let it eat a quarter. Not out of caution. Because the interesting work is on the other side of getting the problem right, and you can't reach it while you're building the wrong thing. Receipts ## Shipped, assessed, and defended in front of people paid to object. 40 ## AI proposals assessed inside a regulated European enterprise Forty use cases went through the framework and some came out classified as not genuinely AI. It held up in front of security review, legal and a works council. Read → 24,000 ## Users in under six months on a fintech platform A utility brief redirected into a discovery product, then built and shipped on the original timeline. It reached 24,000 users in under six months. Read → $1M+ ## Follow-on work from the Jurisage engagement, with AltaML A working AI classifier moved into the workflow researchers already had, instead of asking them to come to it. The engagement led to more than $1M of follow-on work. Read → All three started with someone handing me a brief and asking whether it was the right one. Book a 20-min intro → Several projects, several years, consistently good judgment. I send people to Fanni when they need someone who'll tell them the truth about what they're building. Janis Brix PPC Magic The honest bit ## When you should hire someone else. ## Work with me when - You have a list of AI ideas and no defensible way to compare them - You need someone senior who will say no, and mean it - The governance conversation is coming and nobody owns it - You want the person who assesses it to be the person who builds it ## Hire a firm instead when - You need a team of ten on site from Monday - The work is mainly ML research rather than product decisions - Procurement requires a supplier with bench depth and cover - You want a brand name on the deck more than the answer in it You get seniority on the work rather than on the sales call, a shorter contract, and one person carrying the whole problem. What you don't get is bench depth. If a project needs more hands than I have, I'll say so before we start rather than after. Book a 20-min intro How I work ## Three rules I don't bend. 01 ## Fixed scope before fixed opinions Every engagement starts with a defined deliverable and a price attached. The Reality Check is one week, €1,500–3,500. You know what you're getting before you commit. 02 ## Reasoning written beside every judgement A score you can audit survives a sceptical stakeholder. A score you can't is an opinion with a number attached. Everything I hand over is written so someone outside the project can follow how I got there — and argue with it. 03 ## I don't do the legal work I feed the product and technical assessment into your governance and legal process. I'm careful about that boundary and I say so in the room. Practical ## The things procurement asks. Contracting Happy to sign your NDA before the first call. Invoicing From an EU entity, reverse-charge VAT where applicable. Working with Startups shipping a first AI feature, and enterprises running AI across an organisation. References Available on request, from named clients. Questions ## What people ask before booking. Q ## Who is Fanni Csincsák? An independent AI product consultant based in Budapest, working remotely across the EU, UK and worldwide. Twelve years shipping product in fintech, SaaS and regulated enterprise, including AI governance work inside a regulated European enterprise, and product roles touching Deloitte, ExxonMobil, Novo Nordisk and AltaML. Q ## What does an AI product consultant actually do? Decides which AI ideas are worth building before the budget commits, then builds the ones that are. It is a product job rather than a machine-learning job: the questions that kill most AI projects are about users, workflows and whether a rules engine would have done it better. Q ## Is one person enough for enterprise work? For assessment and product direction, yes — that is exactly where a single senior operator beats a team, because the judgement stays in one head. For delivery at scale it depends on the build. If a project needs more hands than I have, I say so before we start rather than after. Q ## How do engagements start? With a 20-minute call, no charge and no pitch. If it fits, the usual first step is an AI Reality Check — one week, fixed scope, €1,500–3,500. I am happy to sign an NDA before that first call. Your turn ## What's on your roadmap that nobody's willing to question? Twenty minutes, no pitch. Bring the list. Book a 20-min intro See the work Book a 20-min intro ## Analytics? Nothing runs unless you say yes — including session recordings, with typed text masked. Details Strictly necessary Needed for the site to work. Always on, no cookies set. Analytics Google Analytics and PostHog — which pages are read, how long for, where people leave. Session recording PostHog replays a session as a video of the page. All text inputs are masked. Kept 30 days. Accept Reject Choose ====================================================================== # https://www.olively.io/consulting ====================================================================== Skip to content Consulting ## Three ways to work together, with the prices attached . Most AI consulting quotes arrive after two discovery calls and a scoping document. Mine are on this page. You can work out whether this is worth a conversation before you have one. Currently taking on new work. A Reality Check can usually start within two weeks. Book a 20-min intro See the work first Twelve years across - Deloitte - ExxonMobil - Novo Nordisk - AltaML - Cence Where you probably are ## One of these is usually true when someone calls me. ## “We have twelve AI ideas and budget for two.” Every one has a sponsor. None of them are comparable, because each was sized by the person who proposed it. You need a way to rank them that survives the meeting. ## “The board wants an AI strategy by the quarter.” What exists is a list of features with AI written next to them. You need something defensible, and you need to know which items on that list won't survive scrutiny. ## “We built it and nobody uses it.” The model works. Adoption didn't happen, because reaching it meant leaving the tools people already work in. It was run as a technology project when it needed to change how the work happens. That's a product problem, not a model problem. ## “We can't tell whether the pilot worked.” It shipped, people used it for a while, and nobody agreed what success looked like before it started. With no baseline there's nothing to compare against, so it neither succeeds nor dies — it just sits there consuming attention. ## “Nobody here will say this isn't AI.” Someone suspects a project is a rules engine with a new label. Saying so costs them something internally. An outsider can say it for free. The offers ## Start with the diagnostic. The rest follows from what it finds. Most of this budget is already being spent badly. MIT's Project NANDA found roughly 95% of enterprise AI pilots produce no measurable effect on profit and loss — usually not because the technology failed, but because nobody decided up front what success looked like or who owned it. Start here ## AI Reality Check You end the week knowing which AI idea is worth a quarter, which one to kill, and how you'll know either way. €1,500 Startup · up to ten use cases €3,500 Enterprise & mid-market · up to twenty For teams with more AI ideas than budget, and no defensible way to rank them. - Every idea scored against the same six questions - A sequenced build order you can defend to a board - A kill list, with the reasoning beside each call - A baseline per surviving use case — the number today, and what success would have to beat - Where each use case sits under the EU AI Act 40 use cases assessed this way inside a regulated European enterprise — a number of them wasn't AI → Book a 20-min intro Then, if it's real ## AI Build The thing gets built, and survives security and governance review instead of stalling in it. €850–1,100 per day, or €6,000–12,000 fixed scope For teams who know what they're building and need it shipped by someone who has done it inside a regulated environment. - I make the product calls, not just the delivery ones - Built inside your tooling and access controls - Documented so your team owns it after I leave A fintech platform built this way reached 24,000 users in under six months → Talk it through Or, ongoing ## Fractional AI Product Lead AI stops being everyone's side project and becomes one person's responsibility. €6,000–8,000 per month · about a day a week For organisations running AI across several teams with nobody senior owning the whole thing. - I own the roadmap and the trade-offs on it - In your standups, your reviews, your governance forums - Monthly rolling — no minimum term The same judgement turned a stalled AI classifier into $1M+ of follow-on work → Talk it through In their words ## What people say once they've worked with me. I've worked with a lot of product people in AI and R&D. Fanni is the one who asks whether the thing should be built at all. She'd rather kill a bad idea in week one than let it eat a quarter. Csaba Molnar AI Robotics She delivered more than we asked for, and understood the problem better than the brief did. András Zollai Novo Nordisk Working with Fanni was straightforward in a way large-organisation projects usually aren't. She asked the awkward questions early, which meant we didn't discover them late. Akos Hulej ExxonMobil Several projects, several years, consistently good judgment. I send people to Fanni when they need someone who'll tell them the truth about what they're building. Janis Brix PPC Magic Which one ## Almost everyone starts in the same place. ## Start with the Reality Check if You have more AI ideas than you can fund, a board asking what the strategy is, or a suspicion that one of the things on the roadmap isn't really AI. One week and a fixed price tells you which, before you commit a quarter of engineering time. It's also the cheapest way to find out whether we work well together. ## Skip straight to Build or Fractional if You already know what you're building and why, and the gap is capacity or senior product ownership rather than direction. If you can already tell me who uses the thing and what breaks if it doesn't exist, you don't need the diagnostic. I'll tell you on the intro call if I think you're buying the wrong one. Terms ## What this is, and what it isn't. No vendor pitch I sell no platform and take no referral fees. If the answer is a tool you already own, that's the answer. Not legal advice I feed the product and technical assessment into your governance and legal process. I say so in the room. No juniors You work with me directly. There's no team to be handed to after the kickoff call. Contracting NDA before the first call. Invoicing from an EU entity, reverse-charge VAT where applicable. Getting started ## Four steps, and you can stop at any of them. 01 ## A 20-minute call You describe the problem. I tell you which of the three fits, or that none of them do. Happy to sign an NDA before this call rather than after. No charge, and no follow-up sequence if you decide against it. 02 ## A written scope with the price on it Short, specific, and fixed rather than an estimate that drifts. It says what you get, when, and what it costs. If the scope is wrong, we change it before anyone signs anything. 03 ## Intake, then the working session A questionnaire you fill in at your own pace, then one call of about ninety minutes with the people who know the use cases. That's the whole demand on your team's time for a Reality Check. 04 ## Delivery, and a live readout You get the written work, and I present it to the people who have to act on it and answer for every call in it. What happens next is your decision — plenty of teams take the report and build it themselves. Questions ## The ones that come up before signing. How much does AI consulting cost? A one-week AI Reality Check is €1,500 for startups and €3,500 for enterprise and mid-market. Build work is €850–1,100 per day, or €6,000–12,000 fixed scope once the shape is known. A fractional AI product lead is €6,000–8,000 per month for about a day a week. Every price on this page is the price — there is no quote stage. Which one should I start with? Almost everyone starts with the Reality Check. One week, fixed price, and it tells you what's worth building before you commit a quarter of engineering time. It's also the cheapest way to find out whether we work well together. Skip it only if you already know what you're building and why — then the gap is capacity or ownership, and Build or Fractional is the right door. What actually happens during the Reality Check week? You fill in an intake questionnaire listing the use cases. We have one working call of about ninety minutes. At the enterprise tier I also interview stakeholders. Then I score everything, write the reasoning beside each score, and deliver a written report plus a live readout. You are not managing me during that week. What if the assessment says build nothing? That's a legitimate outcome and often the most valuable one. In a forty-use-case assessment inside a regulated European enterprise, several proposals turned out not to be AI at all. Finding that out in week one costs a week rather than a quarter. I'd rather tell you that than take the build fee. What happens after the Reality Check — am I committed to anything? No. It's a fixed-scope engagement that ends with a deliverable. Plenty of teams take the report and build it themselves, which is a fine outcome. If you want me to build it, we scope that separately once we both know what's real. Can you work alongside our existing agency or vendors? Yes, and it's common. I sell no platform and take no referral fees, so I have no reason to recommend replacing a tool that works. If the answer to a use case is something you already own, that's the answer. Do you sign NDAs and work with procurement? Yes. NDA before the first call as standard. Invoicing from an EU entity with reverse-charge VAT where applicable. I've worked inside enterprise security review and works council processes, so the questions are familiar. What if you get hit by a bus? Every engagement produces written artifacts — the scoring, the reasoning and the build order. If I disappear, your team carries on from the documents. For longer engagements I name a backup before we start. Where are you, and what timezones do you work in? Budapest, working remotely across the EU, UK and worldwide. Central European Time, with real overlap for UK and US East Coast. On site where it matters. Your turn ## Bring me the list nobody's willing to question. Twenty minutes, no pitch. If I'm not the right person, I'll say so and point you somewhere better. Book a 20-min intro More on the Reality Check Book a 20-min intro ## Analytics? Nothing runs unless you say yes — including session recordings, with typed text masked. Details Strictly necessary Needed for the site to work. Always on, no cookies set. Analytics Google Analytics and PostHog — which pages are read, how long for, where people leave. Session recording PostHog replays a session as a video of the page. All text inputs are masked. Kept 30 days. Accept Reject Choose ====================================================================== # https://www.olively.io/ai-reality-check ====================================================================== Skip to content AI Reality Check ## An honest read on which of your AI ideas will actually ship . You bring the roadmap. I score every idea against the same six questions and give you a build order you can defend. Book a Reality Check One week  /  fixed scope  /  from €1,500 Why me “She delivered more than we asked for, and understood the problem better than the brief did.” András Zollai  /  Novo Nordisk - 40 AI proposals scored inside a regulated European enterprise — some were not actually AI - EU AI Act governance built and defended in front of legal and a works council - Products that shipped — 24,000 users inside six months on one, a seven-figure follow-on on another - Delivery, not just advice — an enterprise agent taken from fifteen minutes to seven per change request, and policy compliance from 72% to 89% Who this is for ## You have five to twenty AI ideas and no way to compare them. They came from a board meeting, a competitor's website, a vendor pitch and an offsite. Everyone has an opinion about which ones matter and none of the opinions agree. You don't need a strategy deck. You need someone with no stake in the answer to tell you which of these are real. What you get ## Six deliverables, all yours whatever happens next. 01 ## A scored table Every idea marked real, theatre or flagged, with the reasoning written beside it so anyone can audit the call. 02 ## Effort and impact Where each surviving idea sits, with the number it should move and where that number stands today — so the sequencing argument is visible rather than asserted. Where the number doesn't exist yet, that's the first finding. 03 ## A build order The three to five worth doing, in the order I'd do them, and what to do first. 04 ## A kill list What to stop, and what to do instead — often a rules engine, an off-the-shelf tool, or nothing. 05 ## Risk and governance Where each use case lands under the EU AI Act, and what that means before you build. 06 ## A live readout I present all of it to the people who have to act on it, and answer for every score. Sample output — illustrative Auto-summarise support tickets REAL Rules engine, relabelled as AI THEATRE Search over your own documents REAL Chatbot reading a FAQ THEATRE Churn prediction model FLAGGED — NO DATA Flagged means the idea is real but blocked until something changes — usually data that doesn't exist yet, or an owner nobody has named. Sample output — effort against impact The sequencing argument becomes visible rather than asserted. Filled dots are real, dashed are theatre, outlined are flagged. How the scoring works ## The same six questions, asked in the same order, every time. Two gates first. Anything legally prohibited stops there. Anything with no realistic path to the data gets parked rather than scored. Then every surviving idea gets the same treatment, so the answer doesn't depend on who pitched it or how senior they are. First — two gates ## Fail either and it never reaches a score. Is it legal? Prohibited under the EU AI Act stops here Is the data reachable? No path to data means parked, not scored Question 1 — scored 0–3 ## Does this need AI at all? Would a rule, a query or a lookup table do this better? If so, that's a recommendation worth paying for. Boring technology that works beats a model that impresses. Question 2 — scored 0–3 ## Does the data exist? WEIGHTED ×1.5 Not “could we collect it” — does it exist today, at volume, clean enough, with the right to use it. For generative use cases the question changes: is there a test set and governed context, rather than training labels. Question 3 — scored 0–3 ## What's the actual upside? WEIGHTED ×1.5 In euros or hours, against a documented baseline, as a range rather than a single confident number. If nobody can size it, that tells you something. Question 4 — scored 0–3 ## What breaks if it's wrong? Wrong answers are guaranteed. If a wrong answer moves money, breaks a regulation or harms a customer, the bar moves and the governance work comes first. Question 5 — scored 0–3 ## Build or buy? Plenty of genuine AI is a commodity now. Building it yourself is a choice that needs defending — and it changes your legal position. Under the EU AI Act, building can make you a provider; buying does not automatically leave you a deployer, because rebranding a system, modifying it substantially, or changing its intended purpose can move you into provider obligations under Article 25. Question 6 — scored 0–3 ## Who owns it after launch? Models drift and quietly get worse. If nobody owns the monitoring and the retraining, you're not shipping a feature, you're shipping a liability. Finally — one verdict ## Scores resolve to one of three. REAL Worth building. Enters the build order. THEATRE Kill it. Usually a rule, a tool, or nothing. FLAGGED Real, but blocked until something changes. ## The rubric is the easy part. Six questions you could ask yourselves. Most teams don't, and when they do the answers come back wrong — not because the questions are hard, but because of who's answering them. The person who proposed the use case is scoring their own use case. The executive who announced the AI programme is sitting in the room where you decide it's theatre. The data team says the data exists, because saying otherwise means admitting a migration didn't land. None of these people are lying. They just can't give you a clean answer about their own work. I have no stake in which way it goes. That's most of what you're paying for. Where the money goes Most AI budget lands in sales and marketing, where the measured return is lowest. Scores tell you what's real. They don't tell you what to fund first. Surviving contact with the work Most pilots die between demo and production. Judging which will hold up inside a real workflow is a different call from scoring the idea. What it costs to say it Some theatre is expensive to name out loud, especially when someone senior announced it. Knowing when to say it anyway, and how, is judgement. The rubric is calibrated against my own engagements and against the NIST AI risk framework and the EU AI Act. I'll tell you which parts are established practice and which are my judgment — you should know the difference. What it costs ## Fixed price, so the scope can't drift on you. What you are buying How long One week — not a three-month strategy engagement Your time One 90-minute call. I do the rest without you in the room. What lands on your desk A scored table, a build order and a kill list — yours to keep either way What this is not Not a strategy deck. Not an implementation quote. I am not bidding for the build. Startup €1,500 Up to ten use cases - Intake questionnaire - One working call, about ninety minutes - Written report and build order - Live readout Book the startup tier Enterprise & mid-market €3,500 Up to twenty use cases - Everything in the startup tier - Stakeholder interviews - Deeper governance and EU AI Act section - Readout built for leadership Book the enterprise tier “I've worked with a lot of product people in AI and R&D. Fanni is the one who asks whether the thing should be built at all. She'd rather kill a bad idea in week one than let it eat a quarter.” Csaba Molnar  /  AI Robotics “Working with Fanni was straightforward in a way large-organisation projects usually aren't. She asked the awkward questions early, which meant we didn't discover them late.” Akos Hulej  /  ExxonMobil Where it comes from ## Built where saying no had consequences. 40 AI proposals scored some were not actually AI I built this rubric inside a regulated European enterprise and ran forty AI proposals through it. Every score had to survive security review, legal and a works council. Saying no there had consequences. That is where the six questions come from. Most AI assessment frameworks are written by people selling AI. This one was built by someone who had to say no inside an organisation that had already announced it was saying yes. Read that case study → What happens after ## Three options, and I don't mind which you pick. Option 01 ## Take it and go You have the roadmap, the reasoning and the order. Your team builds it. Plenty of clients do this and it's a good outcome. No further cost → Option 02 ## I build the first one Scoped, built, handed over with monitoring in place so it doesn't quietly rot after launch. €850–1,100 / day · €6–12K fixed → Option 03 ## I own the roadmap A fractional arrangement, about a day a week, embedded in your team and in the rooms where AI stalls. €6,000–8,000 / month → The fee isn't contingent on you picking two or three. A diagnostic that only pays if it recommends more work isn't a diagnostic. Questions ## Before you book. ## Why is this paid? Because it takes real work and the answer has value on its own. The kill list alone usually saves more than the fee. If you want a free directional read first, the scorecard takes four minutes. ## What if you tell us to build nothing? Then I've saved you a quarter. That's a legitimate result, and it happens. ## We're not in the EU. Does the AI Act section matter? If you sell into the EU it applies to you. If you don't, I'll cut that section and go deeper elsewhere. ## How much of our time does it take? An intake questionnaire, one call of about ninety minutes, and an hour for the readout. ## Can you sign an NDA? Yes, as standard. ## What if we have more than twenty use cases? Then we scope it as two rounds, or narrow to the ones with budget attached. I'll tell you which on the intro call. Your turn ## Bring me the list nobody's willing to question. Twenty minutes, no pitch. If I'm not the right person, I'll say so and point you somewhere better. Book a Reality Check Try the free scorecard first Book a Reality Check ## Analytics? Nothing runs unless you say yes — including session recordings, with typed text masked. Details Strictly necessary Needed for the site to work. Always on, no cookies set. Analytics Google Analytics and PostHog — which pages are read, how long for, where people leave. Session recording PostHog replays a session as a video of the page. All text inputs are masked. Kept 30 days. Accept Reject Choose ====================================================================== # https://www.olively.io/work ====================================================================== Work ## Six projects, and what each one actually produced. Clients aren’t named. Several were under NDA, and in every case the method is the transferable part rather than the logo. Fintech · consumer product ## Acquiring 24,000 users for a fintech app A card-vault brief, redirected into a fintech lifestyle platform before the build started. Same team, same deadline, different product. 24,000 users in six months Legal tech · applied AI ## Lifting AI adoption without retraining The classifier already worked. Nobody was reaching for it, because using it meant leaving the documents they were already reading. $1M+ follow-on, model unchanged Enterprise ITSM · applied AI ## Integrating an AI platform into ITSM The integration everyone assumed was correct would have made the queue worse. Severity is not urgency, and mapping one onto the other breaks priority. Three nested redirects Enterprise ITSM · conversational AI ## Cutting change requests 15 to 7 minutes The full scope would have shipped something nobody trusted. Cutting it to one path is what moved both numbers. 15 → 7 min, 72 → 89% Enterprise ITSM · product framework ## Turning a per-client survey into one framework The third time the same survey had been built that year. The redirect cost more up front and stopped the fourth and fifth being built at all. Built once, deployed per tenant AI governance · regulated enterprise ## Scoring 40 AI use cases in a regulated enterprise Proposals arriving from every direction, in different languages, judged mostly on who was asking. A framework that survived security and risk review. 40 scored, some not AI Your turn ## What's on your roadmap that nobody's willing to question? Twenty minutes, no pitch. If I'm not the right person, I'll say so and point you somewhere better. Book a 20-min intro Try the free scorecard Archive ## Earlier work, from when this was a product design practice. Eighteen projects from before Olively became an AI consultancy. They are kept because people still find them, and because the reasoning in them holds up even where the positioning has moved on. Logistics SaaS Inventory App for Tracking Stock & Sales Fintech Fintech Figma Design System Home & commercial service SaaS CRM Mobile App Fintech CSV Parsing Bank Statements Fintech Fintech Visual Identity & Branding Home & commercial service SaaS CRM Webapp (Dashboards) Fintech Designing a Product Hunt Launch Fintech Fintech landing page optimisation Home & commercial service Referral and review management web application Kids Streaming tablet app for kids Healthcare Patient-centric healthcare web experience Fintech Tax automation for American expats Healthcare Checkout conversion, from 0.4% to 9% Communication Making remote connections meaningful AI Copy.ai Workflows: a homepage for a Gen AI platform Fintech Column Tax: embedded tax experiences ML Landing page for a machine learning startup Insurance Sola: bridging the gap, one disaster at a time ====================================================================== # https://www.olively.io/work/ai-governance ====================================================================== Case study ## Forty AI use cases. Some of them weren’t AI. The assessment framework behind the Reality Check, built inside an organisation where getting it wrong has consequences. AI governance Regulated enterprise Client not named 40 use cases assessed and classified some filtered out as not genuinely AI 509 hrs a month identified across the quantified cases ## The situation A large regulated European enterprise had AI proposals arriving from every direction — business units, vendors, and internal teams who had all been told to have an AI plan. Nobody could compare them. Each proposal was written in its own language, sized by its own sponsor, and judged mostly on who was asking. Some were substantial. Some were existing systems with a new label. The organisation had no consistent way to tell the difference, and in a regulated environment the cost of getting that wrong isn't just wasted budget. I can't name the client. What I can describe is the method, because the method is the transferable part. ## What happened I built an assessment framework and ran forty use cases through it. Every proposal got the same questions in the same order, scored the same way, with the reasoning written down beside each score. That last part matters more than it sounds: a score you can audit survives a conversation with a sceptical stakeholder, and a score you can't is just an opinion with a number attached. The framework had to hold up in rooms that don't accept hand-waving — security review, legal, and a works council. It also had to sit inside a live regulatory boundary, which meant classification wasn't optional and “we'll deal with compliance later” was never an available answer. I fed the product and technical assessment into that process. I didn't do the legal work, and I'm careful to say so. ## The result Forty use cases assessed and classified. several filtered out as not genuinely AI — proposals a rules engine, an existing system or a query already handled. Around 509 hours a month of identified saving across the two dozen use cases that were quantified. The rejections were the valuable part. Each one was a project that didn't consume a quarter of engineering time to arrive somewhere the organisation already was. What this means if you're reading it This is where the Reality Check comes from. The rubric I use with clients is the same instrument, adapted from an environment where every score had to be defensible to someone whose job was to challenge it. Most AI assessment frameworks are written by people selling AI. This one was built by someone who had to say no inside an organisation that had already announced it was saying yes. What was hard Scoring somebody's proposal low is scoring somebody's budget, headcount and internal case low. In an organisation where every unit had been told to have an AI plan, a framework that filters proposals out is a framework that makes enemies unless the reasoning survives contact with the person it disadvantages. That is why the written reasoning mattered more than the score. A number on its own invites an argument about the number. A number with the reasoning beside it moves the conversation to the reasoning, which is where it should have been. The other difficulty was the ones that were nearly AI. Filtering out an obvious rebrand is easy. Deciding about a proposal that would genuinely benefit from a model but would benefit more from fixing the process underneath it is the judgment the framework had to support rather than replace. ## Common questions ## How do you tell an AI project from a project with AI written on it? Ask whether a rule, a database query or a lookup table would do the job better — not merely adequately. Across forty proposals, a number failed that test: existing systems with a new label. ## What makes an assessment framework hold up? The reasoning written down beside each score. A score you can audit survives a conversation with a sceptical stakeholder; a score you cannot is an opinion with a number attached. ## Why score every proposal the same way? Because otherwise they are judged on who is asking. Each proposal arrives written in its own language and sized by its own sponsor, which makes comparison impossible until the same questions are asked in the same order. Your turn ## Bring me the list nobody's willing to question. Twenty minutes, no pitch. If I'm not the right person, I'll say so and point you somewhere better. Book a 20-min intro See the other projects ====================================================================== # https://www.olively.io/work/cence ====================================================================== Case study ## A fintech utility brief, redirected into a product people actually used. Cence came to me with a reasonable brief. It was also a brief for a product nobody was waiting for. Fintech Consumer product AI Build 24,000 users in under six months 2 app stores, live 0 extra weeks on the timeline ## The situation Cence had funding, a deadline and a spec. What they didn't have was evidence that the spec was the right one. The brief was for a fintech utility — competent, buildable, and aimed at a category where nobody was looking for another entrant. It solved a problem users didn't experience as urgent. ## What happened I pushed back on the brief before building to it. We went back to what the underlying behaviour actually was: what people did before they reached for a tool, and where the friction genuinely sat. That reframed the product from a utility into something closer to discovery — a place people came to find things, rather than a place they came to complete a task. Same team, same funding, same deadline. Different product. Then I built it. Product direction, scope, and the delivery decisions that keep a launch on schedule when the timeline is fixed and the surface keeps wanting to grow. ## The result 24,000 users in under six months , live on the App Store and Google Play. The version in the original brief would have shipped on time too. It would have shipped to a much smaller audience. What this means if you're reading it The most expensive decision on most projects is made before anyone writes code: what gets built. Teams rarely lose a quarter to bad execution. They lose it to executing a brief nobody challenged. What I did for Cence is what the Reality Check does formally — take what a team has decided to build and test whether it survives contact with reality before the money goes in. What was hard Telling a funded team with a fixed deadline that the thing they hired you to build is the wrong thing is not a comfortable conversation, and it is not obviously your job. The spec was not incompetent. It was competent and aimed at the wrong problem, which is harder to argue with than a bad spec. What made it possible was doing the work before the conversation. I did not arrive with an opinion; I arrived with what people actually did before they reached for a tool. That turns a disagreement about taste into a disagreement about evidence, which is a conversation that can be won. The timeline was the other constraint. The redirect had to cost nothing in schedule, or it would have been overruled regardless of whether it was right. ## Common questions ## Why redirect a brief instead of building what was asked for? Because the brief solved a problem users did not experience as urgent. The category already had entrants and nobody was looking for another one. Building it competently would still have produced a product with a small audience. ## Did the redirect cost time? No. Same team, same funding, same deadline — the product changed, the schedule did not. The redirect happened before code was written, which is the only point at which it is cheap. ## What does redirecting a brief actually involve? Going back to the underlying behaviour: what people did before they reached for a tool, and where the friction genuinely sat. In this case that reframed a utility into something closer to discovery — a place people came to find things rather than to complete a task. Your turn ## Bring me the list nobody's willing to question. Twenty minutes, no pitch. If I'm not the right person, I'll say so and point you somewhere better. Book a 20-min intro See the other projects ====================================================================== # https://www.olively.io/contact ====================================================================== Skip to content Contact ## Tell me what you’re trying to build . Or what you’ve been asked to build and aren’t sure about. Two sentences is plenty — enough to tell whether I’m the right person, and I’ll say so either way. Reply within two working days  /  no pitch  / no mailing list ## Send a note Lands in my inbox. I reply to everything within two working days. Your name Email What’s on your mind? Company — leave this empty Send it to me I only use your details to reply. Nothing is added to a mailing list and there is no mailing list. The form is handled by Formspree, who pass it to me — what happens to your data . Other ways ## Three routes, depending on how you work. 01 ## Email If you would rather write it yourself, or attach something. Same inbox, same two days. hello@olively.io → 02 ## Book twenty minutes Better if you want to talk a roadmap through rather than describe it in writing. No pitch, and if I am not the right person I will say so on the call. Pick a slot → 03 ## Score it first Fourteen questions on whether you are ready to build. Four minutes, no email, and the result is often the thing worth sending me. Take the scorecard → The practical bits ## What happens after you send it. - I read everything and reply within two working days. If it is clearly not something I can help with, I will say that rather than leave you waiting. - No pitch on a first call. Twenty minutes, you describe the problem, I tell you whether a Reality Check is the right shape for it or whether you need something else. - Based in Budapest, working across the EU and UK. Remote by default; on-site where a workshop genuinely needs a room. - Under NDA where the work requires it. Most of what I do sits inside regulated organisations, so this is normal rather than an exception. Questions ## The ones that come up before a first conversation. How quickly will you reply? Within two working days, and usually sooner. If you want something faster than that, book the twenty-minute call instead — the calendar shows real availability rather than a request queue. Who would I be contracting with? Olively Korlátolt Felelősségű Társaság (Olively Kft.), a limited liability company registered in Hungary with EU VAT number HU32055164. Full registered details are on the imprint page. Invoices are issued in euros from Hungary, and reverse-charge VAT applies for EU business customers outside Hungary. Will you sign an NDA? Yes, and it is normal rather than an exception — most of this work happens inside regulated organisations. Send yours and I will sign it before we discuss anything specific. I do not require one of my own to have a first conversation. Can you get through our security or procurement review? Usually. The work involves no software installed on your systems and no access to production data. What I typically need is documents, a shared folder and time with people. If your process needs a supplier questionnaire or proof of insurance, ask early and I will fill it in. What happens to what I send you? The contact form is handled by Formspree, who pass it to my inbox and keep a copy for 30 days. Emails and enquiries are kept for two years unless they become a project. Nothing is added to a mailing list, because there is no mailing list. The privacy page has the detail. What if I do not know what I need yet? That is the normal starting point and a good reason to book the call rather than write. If a Reality Check is not the right shape for your problem I will say so, and point you at whatever is — including nobody, if the honest answer is that you should build it yourselves. Do you work outside the EU? Yes. The company is Hungarian and most clients are in the EU and UK, but the work is remote and the time zone is usually the only real constraint. US engagements run through Olively as well. Or start here ## Bring me the list nobody’s willing to question. An AI Reality Check scores every idea on your roadmap and hands back three lists — what is real, what is theatre, and what is blocked. One week, fixed price. How the Reality Check works Book a 20-min intro ## Analytics? Nothing runs unless you say yes — including session recordings, with typed text masked. Details Strictly necessary Needed for the site to work. Always on, no cookies set. Analytics Google Analytics and PostHog — which pages are read, how long for, where people leave. Session recording PostHog replays a session as a video of the page. All text inputs are masked. Kept 30 days. Accept Reject Choose ====================================================================== # https://www.olively.io/use-cases ====================================================================== Use case assessments ## Is it actually AI, and is it worth building? One assessment per use case, with a verdict. Some of these say no, which is the point — scoring forty AI proposals inside a regulated enterprise, a number of them turned out not to be AI at all. Enterprise ITSM · support operations ## AI ticket routing Most ticket routing is a classification problem with a small, stable set of categories and an existing routing table. That is what rules are for. Usually automation with a label Finance operations · accounts payable ## AI invoice matching Matching an invoice to a purchase order and a receipt is deterministic arithmetic. Reading a scanned invoice from a supplier who changes their layout is not. Split — extraction yes, matching no HR · recruitment ## AI CV screening This is real AI and it is explicitly Annex III high-risk under the EU AI Act. The build is the easy part; the obligations attached to it are not. Genuinely AI, and high-risk Internal productivity · collaboration ## AI meeting summaries Real AI, well solved by existing products, and the interesting question is not whether to build but who is being recorded and whether they agreed. Genuinely AI — but buy it Customer operations · support ## AI customer support chatbot Real AI with real value. The failure mode is almost never the model; it is a scope that covers everything and is trusted on nothing. Genuinely AI — scope decides it Legal operations · procurement ## AI contract review Finding clauses and comparing them to a playbook is real AI doing real work. Deciding whether to accept a clause is not a job to hand over. Genuinely AI — with a human gate Supply chain · planning ## AI demand forecasting This is a well-established statistical problem. Classical methods frequently match or beat machine learning on business time series, and they are far cheaper to run. Not new AI — it's forecasting Customer success · subscription business ## AI churn prediction Predicting churn is a solved classification problem. Most projects fail after the prediction, because nobody defined what happens when the score goes red. Genuinely AI — but the model isn't the gap Operations · records management ## AI document classification Twelve stable categories is a rules problem. Two hundred shifting ones is not. The assessment is mostly counting. Depends on the category count Sales · marketing operations ## AI lead scoring Most lead scoring runs on company size, industry and behaviour, all of which are rules. The model appears when you have enough closed-won history to learn from, and most companies do not. Usually automation with a label Finance operations · internal ## AI expense approval An expense policy is a written set of thresholds and categories. Encoding it is automation, and it should stay auditable because employees will contest decisions. Not AI — your policy is already rules Internal productivity · knowledge management ## AI knowledge search Real AI, immediate value, low regulatory surface. The thing that stops it is almost always access control rather than retrieval quality. Genuinely AI — usually the best first project Engineering · developer productivity ## AI code review Real value on a narrow band of review work. Point it at everything and developers stop reading it, which costs you more than not having it. Genuinely AI — narrow it Engineering · quality assurance ## AI test generation Real value on coverage and boilerplate. The trap is that generated tests assert current behaviour, which means they can lock in a bug and defend it. Genuinely AI — with one serious caveat Content operations · localisation ## AI translation Real AI, mature products, and no differentiation available from building your own. The decision is which content gets human review, not which model. Genuinely AI — buy it Fintech · payments · risk ## AI fraud detection Real AI on a genuine pattern problem. What decides whether it works is not detection rate but what happens to the legitimate customers it stops. Genuinely AI — the false positives are the product Operations · shared inboxes ## AI email triage Sender domain, subject patterns and recipient address carry most of the routing signal. That is a rules problem, and rules explain their decisions. Usually automation with a label HR · people operations ## AI onboarding assistant New joiners ask the same questions as everyone else, earlier and more often. A separate system duplicates infrastructure and splits the content maintenance. This is knowledge search — build that Sales operations · revenue ## AI sales forecasting Sales forecasts run on rep-entered stage and close date. If those are optimistic or inconsistent, a model learns the optimism and states it with more authority. Not a model problem — a data problem HR · recruitment operations ## AI resume parsing Extracting structured fields from a CV is ordinary extraction. Ranking candidates is Annex III high-risk. The same product usually does both, and the toggle is not obvious. Genuinely AI — and not the same as screening Customer insight · HR tooling ## AI sentiment analysis Emotion inference in workplace and education contexts is a prohibited practice under the EU AI Act, in force since February 2025. On customer feedback it is permitted and rarely tells you more than reading it would. Prohibited in workplaces — weak elsewhere Content operations · digital asset management ## AI image tagging Real AI, mature and cheap. Unless you are tagging people, in which case you have left ordinary image tagging and entered biometrics. Genuinely AI — buy it Operations · back office ## AI data entry Extraction from documents is real AI and works. But a lot of data entry exists because two systems do not talk to each other, and an integration removes the job rather than automating it. Usually — but ask why it exists Marketing · content operations ## AI content generation Real capability, no differentiation in building it, and an Article 50 transparency duty on published output that has been enforceable since 2 August 2026. Genuinely AI — with a disclosure obligation More assessments are being added. If the one you need is missing, the AI Reality Check scores your whole list in a week. ## Analytics? Nothing runs unless you say yes — including session recordings, with typed text masked. Details Strictly necessary Needed for the site to work. Always on, no cookies set. Analytics Google Analytics and PostHog — which pages are read, how long for, where people leave. Session recording PostHog replays a session as a video of the page. All text inputs are masked. Kept 30 days. Accept Reject Choose ====================================================================== # https://www.olively.io/glossary ====================================================================== Glossary ## The terms, and what people get wrong about them. 33 terms from AI product work and the EU AI Act, each with the mistake that costs most. Definitions are easy to find and mostly identical. What follows each one here is the misreading — the version of the term that gets repeated in meetings and costs money later. Most of these come from assessment work rather than from reading the regulation. The EU AI Act entries are checked against Regulation (EU) 2024/1689 and the European Commission’s implementation timeline as amended by the Digital Omnibus on AI. If you want the dates rather than the vocabulary, what applies from 2 August 2026 covers those. ## EU AI Act ## Provider Under the EU AI Act, the party that develops an AI system or has one developed and places it on the market or puts it into service under its own name or trademark. Providers carry the heaviest obligations, particularly for high-risk systems. What people get wrong That buying rather than building keeps you out of it. Under Article 25 you become a provider if you put your name or trademark on a high-risk system, modify it substantially, or change its intended purpose — and the original provider's obligations transfer to you. ## Deployer The party using an AI system under its own authority, in a professional context. Deployer obligations are lighter than provider obligations but not absent: human oversight, input data relevance, monitoring, and log retention where applicable. What people get wrong That deployer status is permanent. It is a role, not a category, and the same organisation can be a deployer of one system and a provider of another. Rebranding a purchased system moves you. ## Substantial modification A change to an AI system after it has been placed on the market that affects compliance with the Act's requirements, or changes the intended purpose. It can make the modifying party a provider. What people get wrong That fine-tuning a purchased model is routine configuration. Depending on what changes, it can be the thing that transfers provider obligations onto you — and that is a decision usually made by an engineer, not by legal. ## Annex III The list of high-risk use areas in the AI Act: biometrics, critical infrastructure, education, employment, essential public and private services, law enforcement, migration, and administration of justice. Obligations for Annex III systems apply from 2 December 2027 following the Digital Omnibus amendments. What people get wrong That the delay applies to everything. It applies to the high-risk tier only. If you are not doing one of the Annex III activities, December 2027 was never your date. ## Article 50 The transparency provision. It requires disclosure when a person is interacting with an AI system, machine-readable marking of synthetic content, disclosure of deepfakes, and disclosure of AI-generated text published on matters of public interest. Enforceable since 2 August 2026. What people get wrong That it was delayed with everything else. It was not amended by the Digital Omnibus. It is also the part of the Act that reaches the most organisations, because it applies to AI systems generally rather than to a risk classification. ## General-purpose AI model A model displaying significant generality, capable of competently performing a wide range of tasks, that can be integrated into downstream systems. The large language and image models most companies build on. What people get wrong That GPAI obligations started in August 2026. They took legal effect on 2 August 2025. What changed in 2026 is that enforcement of them began. ## AI literacy The Article 4 requirement that providers and deployers take measures to ensure a sufficient level of AI competence among staff dealing with AI systems, proportionate to context and risk. Applicable since 2 February 2025. What people get wrong That it means buying a training course. It means being able to show the people operating a system understand what it does and where it fails. A generic e-learning module with no relation to the systems you actually run satisfies procurement, not the requirement. ## Conformity assessment The process of demonstrating that a high-risk AI system meets the Act's requirements before it goes to market. Depending on the system, this is a self-assessment or involves a notified body. What people get wrong That it happens at the end. Most of what a conformity assessment asks for — risk management, data governance, technical documentation, logging — has to exist while the system is being built. Retrofitting it is the expensive version. ## Post-market monitoring The obligation on providers of high-risk systems to actively collect and review performance data after deployment, and to act on what it shows. What people get wrong That it is a reporting duty. It is an operational one. If nobody owns the system after launch, there is nothing to monitor with and no one to act on what monitoring finds. ## Prohibited practices The uses banned outright: social scoring, exploitative manipulation, untargeted facial image scraping, emotion inference in workplaces and education, and certain biometric categorisation. In force since 2 February 2025. What people get wrong That the list is exotic and irrelevant to ordinary business. Emotion inference in a workplace context catches more HR and productivity tooling than most buyers expect. ## Deciding what to build ## Use case A specific job an AI system is proposed to do, described concretely enough that its value and feasibility can be assessed independently of other proposals. What people get wrong That a theme is a use case. "Improve customer experience with AI" cannot be scored, funded or falsified. If two people would build different things from the same description, it is not yet a use case. ## AI theatre Work that carries an AI label without needing to. Usually automation, a rule, a lookup table or a database query wearing a more fundable name. What people get wrong That it is dishonest. It is almost always sincere. The incentives point that way: calling a project AI gets a faster answer than calling it automation, and nobody in the room is rewarded for saying this one is not AI. ## Kill criteria The condition, agreed before a project starts, under which it will be stopped. Typically a named metric failing to move by a named date. What people get wrong That agreeing them signals a lack of confidence. The opposite: a pilot with no defined way to fail does not end, it becomes permanent — consuming budget and producing nothing while nobody is willing to be the person who cancels it. ## Baseline The current value of the metric a project is meant to move, measured before work starts. What people get wrong That it can be established afterwards. Without a baseline there is no way to demonstrate the change, which is how a project can ship successfully and still be unable to prove it was worth doing. ## Proof of concept A small build intended to establish whether something is technically possible. What people get wrong That it is a step toward production. It answers a feasibility question and is usually built on hand-prepared data with no operational constraints. IDC found that for every 33 proofs of concept an enterprise starts, roughly four reach production. ## Pilot A limited deployment in real conditions, intended to establish whether something delivers value. What people get wrong That it is a longer proof of concept. A pilot without a baseline, a target and a date is not measuring anything — it is a proof of concept with a larger audience. ## Shadow AI AI tools adopted by individuals or teams outside sanctioned procurement, security review or governance. What people get wrong That an inventory of approved tools describes your AI footprint. It describes the part you know about. Under the AI Act your obligations attach to what is actually in use, not to what was approved. ## Building and running ## Retrieval-augmented generation An architecture where a model is given relevant documents retrieved at query time, rather than relying only on what it learned during training. Commonly abbreviated RAG. What people get wrong That it eliminates hallucination. It reduces one cause of it. A model handed the wrong documents produces a confident answer grounded in the wrong source, which is harder to catch than an obvious invention. ## Fine-tuning Further training of an existing model on a specific dataset to adapt its behaviour. What people get wrong That it is the default answer to poor output. It is expensive, it has to be repeated when the base model updates, and prompt design or retrieval solves the problem more often than the industry's enthusiasm for fine-tuning suggests. It can also constitute substantial modification under the AI Act. ## Evaluation The systematic measurement of a model's output quality against a known set of cases. Often shortened to evals. What people get wrong That it is a testing phase. It is a permanent cost. Models change under you, inputs drift, and without ongoing evaluation the first sign of degradation is a user complaint. ## Golden dataset A curated set of inputs with known-correct outputs, used as the reference for evaluation. What people get wrong That it can be assembled by the engineering team alone. The value depends on the correct answers being correct, which usually requires the people who do the job — and their time is the part nobody budgets for. ## Hallucination Output that is fluent, confident and wrong. A model generating plausible content unsupported by its inputs or by fact. What people get wrong That the fluency is a bug. It is the system working as designed — these models optimise for plausible continuation, not for truth. Which is why detection has to be external to the model. ## Guardrails Controls constraining what a system will accept as input or produce as output. What people get wrong That they are a safety layer added at the end. Guardrails written after launch are written in response to incidents, which means each one is documenting something that already went wrong. ## Drift Degradation in performance over time as real-world inputs diverge from what the system was built and tested against. What people get wrong That it announces itself. Drift is gradual and the system keeps producing confident output throughout. Without monitoring, the detection mechanism is a customer noticing. ## Observability The ability to see what a system is actually doing in production: what it was asked, what it returned, how long it took, what it cost, and where it failed. What people get wrong That logging is observability. Logs tell you what happened. Observability means being able to answer a question you had not thought to ask in advance, which is what an incident requires. ## Human in the loop A design in which a person reviews or approves AI output before it takes effect. What people get wrong That naming a reviewer creates oversight. If the reviewer has no time, no context and no incentive to disagree, the loop is decorative — and under the AI Act, human oversight has to be effective rather than nominal. ## Prompt injection An attack where instructions hidden in content the model processes cause it to ignore its original instructions. What people get wrong That it is a prompt engineering problem. If your system reads untrusted content — emails, documents, web pages — it is an architecture problem, and no amount of instruction hardening closes it fully. ## Agent A system that takes actions rather than only producing text: calling tools, writing to systems, or executing multi-step tasks with limited supervision. What people get wrong That agentic is a capability tier. It is a risk tier. The moment a system acts rather than suggests, the cost of being wrong stops being a bad answer and starts being a bad outcome. ## Context window The amount of text a model can consider at once, measured in tokens. What people get wrong That a larger window removes the need for retrieval. Filling a window with everything available degrades output and multiplies cost. What goes in still has to be selected. ## Token The unit of text a model processes, typically a word fragment. Pricing and context limits are counted in tokens. What people get wrong That token cost is the cost. It is the visible part. Evaluation, observability, governance, security review and human quality assurance are usually larger, and they are the line items missing from the business case that got approved. ## The economics ## Run cost The ongoing cost of operating an AI system in production: inference, evaluation, observability, governance, model updates, drift management, security review and human quality assurance. What people get wrong That it is the API bill. When the real run cost surfaces, the return in the original business case does not evaporate — it was never there. The case contained an arithmetic error from the start. ## Time to value How long from starting work to the point where a measurable outcome changes. What people get wrong That shortening it means moving faster. It usually means scoping smaller. A narrow thing in production beats a broad thing in development, and the difference is a scoping decision rather than a delivery one. ## Build versus buy The decision between developing an AI capability internally and licensing an existing one. What people get wrong That building gives you control and buying gives you speed. Both can be true and neither is automatic — and the decision has legal consequences under the AI Act, because building can make you a provider while buying does not reliably keep you a deployer. Seen in practice Scoring 40 AI use cases in a regulated enterprise — where most of these misreadings turned up, and what they cost. ## Analytics? Nothing runs unless you say yes — including session recordings, with typed text masked. Details Strictly necessary Needed for the site to work. Always on, no cookies set. Analytics Google Analytics and PostHog — which pages are read, how long for, where people leave. Session recording PostHog replays a session as a video of the page. All text inputs are masked. Kept 30 days. Accept Reject Choose ====================================================================== # https://www.olively.io/scorecard ====================================================================== Skip to content AI readiness scorecard ## Fourteen questions on whether you’re ready to build AI . Most AI work fails on things you could have checked first — a problem nobody can state without saying “AI”, data that isn’t there yet, a payoff nobody sized. This scores the five that matter and names the one most likely to stall you. 4 minutes  /  14 questions  /  no email required Question 1 of 14 ## Back Back to my result Saved in this browser as you go, so a refresh doesn’t lose your place. Your score — ## Weakest area This scores how ready you are in general. It doesn’t score your individual ideas — for that, every one has to go through the full rubric on its own. That’s the Reality Check. See what a Reality Check covers Send me your score Sending it opens your own mail app with the answers already filled in, and I’ll write back with where I’d start. If nothing opens, your browser has no mail app set — use Copy result below instead. Need to show someone else? Copy it, or print this page to PDF. Copy result Print Review my answers Start again Who this is for ## You have AI ideas and no shared way to judge them. They arrived from a board meeting, a vendor pitch, a competitor’s website and somebody’s offsite. Everyone has a view on which ones matter, and the views don’t agree. Before you can argue about which idea to build, you need to know whether you are in a position to build any of them well. What gets scored ## Five things, and none of them are about the technology. Every question belongs to one of these. The score you get is the average, and the area that comes out lowest is the one most likely to stall you. 01 ## What the AI is for Whether the problem can be described without the word AI, and whether a rule or a query would do it better. Most of what fails, fails here — the questions that catch it . 02 ## Data, and knowing it works Data readiness in the literal sense — whether it exists today rather than could be collected — and whether anyone has decided how you would tell a good output from a bad one against a baseline. 03 ## The payoff Whether the value has been sized in hours or euros against something you already measure. Unsized work is the first thing cut when budgets tighten — see which numbers hold up . 04 ## Risk and sign-off What happens when the model is wrong, and whether anyone has checked where this sits under the EU AI Act. A low score here means AI governance comes first, never don’t build — how that ran in practice . 05 ## Who runs it after launch Whether a named person owns it, and whether the way people work is expected to change. This is where pilots die between demo and production — models drift, and unowned ones quietly stop being used. What you get ## A number, five bars, and the one area to fix first. 01 ## A score out of 100 With a band that says plainly what it means — from high theatre risk through to ready. Most roadmaps land in the middle before anyone has interrogated them. 02 ## A breakdown by area All five scored separately, so a strong average doesn’t hide one area that will stop the work. 03 ## Your weakest area, named With a plain explanation of what typically goes wrong there and what it costs when nobody catches it early. Answers that don’t apply to you are excluded from the score rather than marked zero — a twelve-person company isn’t penalised for having no legal department. What happens to your answers ## Saved on your device, and not sent to me. - Your answers are stored in this browser — so a refresh doesn’t lose your place, and a finished score is still there if you come back. They stay until you clear your browser data or press Start again . - The answers themselves never leave your device. There is no account, no server storing them, and nothing arrives in my inbox when you finish. - If you accepted analytics, one event is recorded when you finish: the score, which area came out weakest, and the team-size band you picked — never the individual answers. If you rejected analytics, or your browser sends a Global Privacy Control signal, nothing is recorded at all — the cookie page has the detail. - Nothing is emailed unless you click the button that does it , and that opens your own mail client with the text visible so you can see exactly what you would be sending. Who built this ## The scoring comes out of work that was done, not a framework I read about. “Fanni is the one who asks whether the thing should be built at all.” Csaba Molnar  /  AI Robotics - 40 AI proposals scored inside a regulated European enterprise — some were not actually AI - EU AI Act governance built and defended in front of legal and a works council - 24,000 users in under six months on a fintech platform shipped end to end Questions ## What this scores, and what it doesn’t. What is an AI readiness scorecard? A short structured assessment of whether a team is in a position to build AI successfully — not whether a particular idea is good. It scores five things: whether the problem is defined without reaching for the word AI, whether the data exists today, whether the payoff has been sized, whether the consequences of a wrong answer are understood, and whether anyone will own the thing after launch. How long does the AI readiness scorecard take? About four minutes. Fourteen questions, one screen at a time, and you get your score immediately without giving an email address. Does it score my specific AI idea? No. It scores how ready you are in general — whether the problems are defined, the data exists, the payoff is sized, the risk is understood and someone will own the result. Scoring an individual use case takes the full rubric, one idea at a time. Do I have to give an email address to see my score? No, and there is no email step at any point. The score and the per-area breakdown appear straight away and are yours to copy or print. If you want a view on what to do about the result, you can send it to me and I will write back — but that is a choice, not a gate. What counts as a good score? Most teams land in the middle band before anyone has interrogated their roadmap, so a mixed result is the common case rather than a bad one. A high score usually means the thinking is done and what is needed is delivery rather than diagnosis. The number matters less than which of the five areas came out weakest. Is this the same as an AI maturity assessment? No. Maturity models score an organisation against a ladder of capability — tooling, talent, governance, platform. This scores whether a specific piece of work is ready to start. A company can be mature and still propose something with no data behind it, and a five-person startup can be ready on every question that matters. What does a low score mean? Usually that the thinking hasn't been done yet, not that the idea is bad. Most roadmaps score in the middle band before anyone has interrogated them. A low score on risk in particular means governance comes first — it never means don't build. What happens next ## A score tells you where you stand. It doesn’t tell you what to build. The Reality Check scores every idea on your roadmap one at a time and hands back three lists — what is real, what is theatre, and what is blocked until something changes. One week, fixed price, and I have no stake in which way any of it goes. How the Reality Check works Book a 20-min intro ## Analytics? Nothing runs unless you say yes — including session recordings, with typed text masked. Details Strictly necessary Needed for the site to work. Always on, no cookies set. Analytics Google Analytics and PostHog — which pages are read, how long for, where people leave. Session recording PostHog replays a session as a video of the page. All text inputs are masked. Kept 30 days. Accept Reject Choose