In brief. MIT NANDA measured that 95% of enterprise generative AI projects produce no measurable impact on the P&L. This is not a model problem: it is a problem of integration into processes, of daily monitoring, and of who answers for it. Companies that build with a specialised partner succeed roughly 67% of the time, against a third for purely internal builds. The 5% that works starts from a bounded scope, a problem measured in euros, and one person who follows it every day.
A pilot that worked beautifully, until it answered a real customer
A financial institution with a few thousand employees had just finished rolling out AI agents across the whole organisation. The project had come down from above, sponsored at board level, and the demo had been flawless: fast answers, correct tone, use cases covered one after another. Nobody in the room had anything to object to.
Then somebody noticed an answer that did not add up. Not a spectacular error, the kind of imprecision that goes unnoticed until it reaches the wrong customer at the wrong moment. That is when the question nobody had asked before launch surfaced: who checks, every day, what these agents actually answer? The answer was nobody. The pilot technically worked. The process that was supposed to contain it, monitor it and correct it did not exist.
The pattern has a name, and MIT measured it at scale
What we saw in that case is not an isolated incident. It is exactly the pattern that The GenAI Divide: State of AI in Business, the 2025 research from MIT's NANDA initiative, measured across 300 public deployments, 150 manager interviews and a survey of 350 employees: only 5% of enterprise generative AI projects lead to real acceleration in revenue or margin. The remaining 95% leave no measurable impact on the P&L.
Building alone fails two times out of three.
The figure made noise because it sounds like a verdict on the technology. Read from inside a real project, it says something else: failure is not a rare, unpredictable event, it is almost always the same thing repeating. MIT calls it the learning gap, the distance between a model that answers well in the lab and an organisation that knows what to do with it every day.
Why it is not the models' fault
The common reflex, faced with an AI project that produces no results, is to blame the model: it does not understand enough, it gets too much wrong, it is not mature yet. The research contradicts that reading. General-purpose models work today, and in the lab they behave well nearly all the time. What is missing is not in the model. It is between the model and the actual work.
"The pilot technically worked. The process that was supposed to contain it, monitor it and correct it did not exist."
The bank case shows it well: the technology did not need to be better. It needed a defined scope (which questions it can handle alone, which it cannot), somebody reading a sample of conversations every week, and a clear criterion for when to step in. Without those three things, even the best model in the world produces the same outcome: it works in the demo, it breaks on the first anomalous case, and nobody notices until it is already a problem.
The difference shows in the scope, not in the size
Last year we worked with a construction general contractor of about 50 people that wanted to automate one single process: sending and tracking customer quotes. No all-encompassing agent, no ambition to cover every department. One process, acceptance criteria defined before starting, and one internal person who checked every week that the system did exactly what it was supposed to do. Today that process runs, it measures the time saved, and nobody in the company wonders whether it "works": they see it in the numbers.
The difference between this case and the bank is not the size of the company, nor the budget, nor the sophistication of the model used. It is that in the second case the scope was clear from day one, and somebody answered for it. That is exactly the line between the 5% and the 95% described by MIT, seen up close: the winner is not whoever has the most advanced AI, it is whoever set up better what it has to do, who checks it, and how the result gets measured.
How to set up a project so it lands in the 5%
The way we work at Morfeus exists precisely to avoid the mistake we have seen repeat: starting from the tool instead of the problem, staying in demo, letting the system live without anyone answering for it.
- Diagnosis. Before talking technology, we measure where the company is losing value every day. With the ROIometer that feeling becomes a number, and the points of loss become Value Leaks quantified in euros.
- System. We build on a bounded scope, in production from day one, not in an isolated demo. The system sits on MARF, the infrastructure that stays in the company and extends over time instead of starting from zero with every project.
- Value. Every month a Value Report shows what the system actually produced, in euros, not in slides.
- Autonomy. We train an AI Champion per department: the person who checks, corrects and evolves the system every day, so the capability stays inside the company the day after we leave.
| Cause of failure | The signal inside the company | Countermeasure |
|---|---|---|
| No defined scope | "Let's do AI everywhere", no priority use case | One process, acceptance criteria |
| Nobody answering day to day | The system runs, but it belongs to no one | Departmental AI Champion |
| Zero monitoring of answers | Nobody reads what the system tells customers | Weekly review on a sample |
| Technical metrics, not euros | "92% accuracy" but the CFO cannot see the impact | Monthly Value Report |
| Project handed down by IT | The operating departments were not in the room | Diagnosis with the people who live the process |
The thread running through these four steps is the same one we saw missing in the bank case and present in the general contractor: a clear scope, and somebody taking care of it every day.
In short
The 95% MIT measured is not a verdict on the technology, it is a diagnosis of method. A pilot can be technically flawless and still fail, if nobody integrates it into the processes and nobody answers for it every day. The question to start from is not which AI to buy, but where you are losing value today, and who will be checking it tomorrow.
Want to know which side you are on?
In a few minutes the ROIometer tells you what a process costs you today and how much you can recover. It is the first step to staying out of the 95%.