Insights · Margin & ROI
Margin & ROI · June 29, 2026 · 7 min

Why 95% of AI projects fail (and how to be in the 5% that works)

MIT NANDA measured that 95% of enterprise generative AI projects produce no measurable impact on the P&L (The GenAI Divide, 2025). This is not a model problem: it is a problem of integration into processes, of daily monitoring, and of who answers for the results. Projects built with specialised partners succeed roughly 67% of the time, against a third for purely internal builds. The 5% that works starts from a bounded scope, a problem measured in euros, and one person who follows it every day.

MA
Matteo Arnaboldi
CEO & Co-Founder, Morfeus

Updated on July 9, 2026

MARGIN & ROI
In brief

MIT NANDA measured that 95% of enterprise generative AI projects produce no measurable impact on the P&L (The GenAI Divide, 2025). This is not a model problem: it is a problem of integration into processes, of daily monitoring, and of who answers for the results. Projects built with specialised partners succeed roughly 67% of the time, against a third for purely internal builds. The 5% that works starts from a bounded scope, a problem measured in euros, and one person who follows it every day.

In brief. MIT NANDA measured that 95% of enterprise generative AI projects produce no measurable impact on the P&L. This is not a model problem: it is a problem of integration into processes, of daily monitoring, and of who answers for it. Companies that build with a specialised partner succeed roughly 67% of the time, against a third for purely internal builds. The 5% that works starts from a bounded scope, a problem measured in euros, and one person who follows it every day.

A pilot that worked beautifully, until it answered a real customer

A financial institution with a few thousand employees had just finished rolling out AI agents across the whole organisation. The project had come down from above, sponsored at board level, and the demo had been flawless: fast answers, correct tone, use cases covered one after another. Nobody in the room had anything to object to.

Then somebody noticed an answer that did not add up. Not a spectacular error, the kind of imprecision that goes unnoticed until it reaches the wrong customer at the wrong moment. That is when the question nobody had asked before launch surfaced: who checks, every day, what these agents actually answer? The answer was nobody. The pilot technically worked. The process that was supposed to contain it, monitor it and correct it did not exist.

The pattern has a name, and MIT measured it at scale

What we saw in that case is not an isolated incident. It is exactly the pattern that The GenAI Divide: State of AI in Business, the 2025 research from MIT's NANDA initiative, measured across 300 public deployments, 150 manager interviews and a survey of 350 employees: only 5% of enterprise generative AI projects lead to real acceleration in revenue or margin. The remaining 95% leave no measurable impact on the P&L.

The partner vs in-house gap

Building alone fails two times out of three.

100% 50% 0% 67% 33% With a specialised partner In-house build Success rate of enterprise generative AI projects
Specialised partnerIn-house build
Source: MIT NANDA, The GenAI Divide 2025, across 300 deployments. It is not the model that makes the difference: it is having already seen where the integration breaks.

The figure made noise because it sounds like a verdict on the technology. Read from inside a real project, it says something else: failure is not a rare, unpredictable event, it is almost always the same thing repeating. MIT calls it the learning gap, the distance between a model that answers well in the lab and an organisation that knows what to do with it every day.

Why it is not the models' fault

The common reflex, faced with an AI project that produces no results, is to blame the model: it does not understand enough, it gets too much wrong, it is not mature yet. The research contradicts that reading. General-purpose models work today, and in the lab they behave well nearly all the time. What is missing is not in the model. It is between the model and the actual work.

"The pilot technically worked. The process that was supposed to contain it, monitor it and correct it did not exist."

The bank case shows it well: the technology did not need to be better. It needed a defined scope (which questions it can handle alone, which it cannot), somebody reading a sample of conversations every week, and a clear criterion for when to step in. Without those three things, even the best model in the world produces the same outcome: it works in the demo, it breaks on the first anomalous case, and nobody notices until it is already a problem.

The difference shows in the scope, not in the size

Last year we worked with a construction general contractor of about 50 people that wanted to automate one single process: sending and tracking customer quotes. No all-encompassing agent, no ambition to cover every department. One process, acceptance criteria defined before starting, and one internal person who checked every week that the system did exactly what it was supposed to do. Today that process runs, it measures the time saved, and nobody in the company wonders whether it "works": they see it in the numbers.

The difference between this case and the bank is not the size of the company, nor the budget, nor the sophistication of the model used. It is that in the second case the scope was clear from day one, and somebody answered for it. That is exactly the line between the 5% and the 95% described by MIT, seen up close: the winner is not whoever has the most advanced AI, it is whoever set up better what it has to do, who checks it, and how the result gets measured.

How to set up a project so it lands in the 5%

The way we work at Morfeus exists precisely to avoid the mistake we have seen repeat: starting from the tool instead of the problem, staying in demo, letting the system live without anyone answering for it.

  • Diagnosis. Before talking technology, we measure where the company is losing value every day. With the ROIometer that feeling becomes a number, and the points of loss become Value Leaks quantified in euros.
  • System. We build on a bounded scope, in production from day one, not in an isolated demo. The system sits on MARF, the infrastructure that stays in the company and extends over time instead of starting from zero with every project.
  • Value. Every month a Value Report shows what the system actually produced, in euros, not in slides.
  • Autonomy. We train an AI Champion per department: the person who checks, corrects and evolves the system every day, so the capability stays inside the company the day after we leave.
The five most common causes of failure, and the Morfeus countermeasure
Cause of failureThe signal inside the companyCountermeasure
No defined scope"Let's do AI everywhere", no priority use caseOne process, acceptance criteria
Nobody answering day to dayThe system runs, but it belongs to no oneDepartmental AI Champion
Zero monitoring of answersNobody reads what the system tells customersWeekly review on a sample
Technical metrics, not euros"92% accuracy" but the CFO cannot see the impactMonthly Value Report
Project handed down by ITThe operating departments were not in the roomDiagnosis with the people who live the process

The thread running through these four steps is the same one we saw missing in the bank case and present in the general contractor: a clear scope, and somebody taking care of it every day.

In short

The 95% MIT measured is not a verdict on the technology, it is a diagnosis of method. A pilot can be technically flawless and still fail, if nobody integrates it into the processes and nobody answers for it every day. The question to start from is not which AI to buy, but where you are losing value today, and who will be checking it tomorrow.

Want to know which side you are on?

In a few minutes the ROIometer tells you what a process costs you today and how much you can recover. It is the first step to staying out of the 95%.

Try the ROIometer
Frequently asked

In three answers

What do AI project failures actually come down to?

Not model quality, which today is high. They come down to how far AI gets integrated into the real processes: a clear scope, someone responsible day to day, constant monitoring. Without those three, even the best model stays a demo.

Is it better to build AI in-house or with a partner?

The MIT research shows that solutions built with specialised partners succeed roughly 67% of the time, against a third for purely internal builds. Not because the partner has a better model, but because they have already seen where the integration breaks.

How do I know if my AI project is about to fail?

Typical signals: nobody can say which number is supposed to improve, there is no person checking what the system answers every day, the project has been stuck at pilot stage for months, and its fate is decided by IT alone, without the operating departments.

Measure, before anything

The problem you don't see has a price.

Try the ROIometro: pick a department and see, in euros, where your company loses value every day.