When a client tells me "AI isn't paying off the way it promised", the first thing I ask is not which model they use or how much they spent. I ask: "what did you mean by 'paying off', and when did you decide it?" In most cases the answer is a silence, followed by a number improvised on the spot. Hours saved, maybe. How many people open the tool every day, perhaps. Almost never a euro.
In brief. AI ROI is not measured by counting hours saved or active licences: that is a usage criterion, not a value criterion. It is measured by defining, before you start, a single number in euros to verify every month. McKinsey (State of AI 2025, 1,993 companies) confirms the pattern from the outside: 88% of companies use AI regularly, but only 5.5% report real financial impact. The gap is almost always the same one: nobody decided up front what to count.
The moment it stops being up for discussion
In the projects we have run for years there is a turning point that comes back the same way, client after client. It is not when the system goes live, and it is not even when it works well technically. It is the month the client stops saying "how many hours did it save us" and starts saying "how many euros did it recover". From that moment the project stops, almost always, being a line item up for debate at renewal.
This is not a marketing trick: it is a change of unit. Hours saved are an estimate, usually an optimistic one, that nobody really checks after the first month. Euros recovered are a number you can verify, put next to the P&L and defend in front of a CFO. Whoever moves from the first to the second stops having to "convince" anyone that AI works: they demonstrate it.
That is why, before starting on any system, we define a Value Leak with the client: the precise point where the company is losing margin today, quantified in euros before the tool is even chosen. This is not bureaucracy. It is the only thing that makes it possible, twelve months later, to answer "what did it return" without inventing a number on the spot.
What McKinsey says, and why it confirms the pattern
The State of AI 2025 report by McKinsey, run across 1,993 companies, puts a precise number on something we have been seeing in the field for a while: 88% of companies report regular AI use in at least one function. Only 5.5% report real financial impact, measurable at P&L level.
One company in twenty sees the return. The other nineteen are paying for something they cannot quantify, not because the tool does not work, but because they never defined what would have to happen for them to be able to say so.
Out of 100 companies using AI, 6 report real financial impact.
Looking deeper into the data, only 39% of the companies surveyed report measurable impact at company EBIT level. At single-function level the picture changes: software engineering, manufacturing and IT report cost reductions between 10% and 20%, marketing and product development report revenue increases above 10%. So the ROI almost always exists. It just stays locked inside the department that generated it, because nobody aggregates it into a single number that is comparable over time, the kind of number we would use to judge any other investment in the company.
| Function | Type of impact | Reported range |
|---|---|---|
| Software engineering | Operating cost reduction | 10-20% |
| Manufacturing | Operating cost reduction | 10-20% |
| IT | Operating cost reduction | 10-20% |
| Marketing | Revenue increase | >10% |
| Product development | Revenue increase | >10% |
| Company EBIT (aggregate) | Impact reported at P&L level | 39% of companies |
Why usage is not the same thing as measurement
The most common cause is not model quality, which today is generally high for the kind of work companies ask of it. It is that most organisations measure adoption, not value: how many people use the tool, how many licences are active, how much time is spent interacting with AI. None of those numbers tells you whether the company's margin moved by a single euro.
It is the difference between a usage dashboard and a value statement. The first tells you how much the tool works. The second tells you how much it returned. They are two different questions, and answering the first one well guarantees nothing about the second: you can have very high adoption and a return of zero, which is exactly where the majority of the companies McKinsey surveyed find themselves.
The pilot that never reaches production
There is a second cause, connected to the first: only about a third of the companies that took AI to pilot stage managed to scale it into production. The other two thirds are stuck in an intermediate phase where the system exists, gets shown in a demo that impresses, but never really enters the daily workflow. The same pattern shows up in the MIT research on AI project failure: a pilot applauded in a meeting is not a system that generates value. Those are two different things, and they are easy to confuse.
The companies that get past this point have three things in common, according to the same report: real leadership involvement, a goal tied to a business indicator defined before starting, and a genuine redesign of the process around the new tool rather than a tool bolted on top of the existing flow. Only 21% of companies reach that level of redesign. It is the difference between measuring a return and measuring only usage, the same distinction again.
Leadership inside
Not a formal sponsor: a decision maker who owns the number month after month.
KPI up front
One value criterion in euros, decided before a line of code is written.
Redesigned process
The system goes inside the flow, not on top of it. Only 21% actually get there.
How you actually measure it: one number, not a dashboard
AI ROI is not measured with a dashboard of usage metrics. It is measured with a single value criterion, defined before the project starts rather than reconstructed afterwards to justify it. In practice: how much margin does this system recover every month, against a starting point already quantified as a measured Value Leak.
This is where the Value Report comes from: the monthly statement with which we answer "what did it return", not a list of activities carried out but the value generated in euros, verified month on month. Across the projects we have run since 2023, this criterion has let us account for more than 4 million euros of recovered margin across over 60 systems in production. That is not a slide number: it is the sum of what happens when, before writing a line of code, you decide what you are going to measure.
The change of unit. From a usage dashboard (licences, hours, clicks) to a value statement (euros recovered month on month). Illustrative example on a real baseline.
In short
AI in business today is used almost everywhere and pays off rarely. The difference between ending up in the 5.5% and staying outside it is not the technology chosen: it is having defined, before starting, a single value criterion in euros, verified every month. Whoever measures usage knows how much the tool works. Whoever measures ROI knows how much it returned, and stops having to defend it.
What did your AI project return, in euros?
Try the ROIometer: define in a few minutes the value criterion to verify every month, before you even choose the tool.