Insight · AI adoption
Five structural reasons projects die between the demo that impressed everyone and the system nobody ended up using — and what to fix first in each case.
The numbers are worse than most leadership teams assume. Deloitte's 2026 survey of more than 3,000 executives across 24 countries found that only 25% of companies have moved more than 40% of their AI pilots into production, and 84% have not redesigned any of the underlying work around AI. BCG's parallel research found roughly 5% of organisations generating substantial value while 60% generate none at all.
What makes this interesting is that the gap is not explained by budget or by model quality. Companies with large budgets and excellent models are well represented among the 60%. The failures are structural, and they repeat.
A demo answers "can a model do this?" — a question whose answer has been yes for a while. The question that decides production is different: can it do this reliably enough, at your volume, on your messiest inputs, when the person using it is busy and will not check the output?
Pilots are usually run on clean, curated examples by people who want them to succeed. The gap between that and Tuesday afternoon in a real operations team is where the project quietly dies.
What to do differently: define the production acceptance bar before the pilot starts, in numbers. What accuracy, at what latency, on what proportion of real inputs. Then test against a sample drawn from actual historical work, including the awkward cases.
This is the most common and the least discussed. The pilot was built by a vendor, a consultancy, or one enthusiastic person in the business. When it moves to production, the question of who maintains the prompts, monitors the outputs, updates the documents behind it and answers questions about it has never been settled.
An unowned system degrades. Documents go stale, prompts drift out of alignment with policy, and within two quarters people quietly stop trusting it.
What to do differently: name the internal owner before the build begins, not at handover, and make sure that person is trained during the build rather than briefed after it. If no such person exists, that is a finding worth acting on before writing any code.
Data quality is where mid-market AI projects most often fail silently. A use case that works beautifully against a hand-picked set of documents behaves differently against ten years of inconsistently filed ones. Fields are missing, formats changed twice, and the system of record disagrees with the spreadsheet everyone actually uses.
What to do differently: assess data availability, accuracy and lawful usability as part of use-case selection rather than after it. Some of the highest-value opportunities on a roadmap should be sequenced later precisely because the data groundwork has to come first. That sequencing decision is most of the value of a good readiness assessment.
In regulated sectors this is the wall almost everything hits. The pilot works, the business wants it, and then someone in risk or compliance asks how a decision was reached, what data trained or informed it, where the human sits in the process, and what happens when it is wrong. If those answers have to be reconstructed after the fact, they are usually reconstructed badly, and the project goes back into review indefinitely.
Recognised frameworks — the NIST AI Risk Management Framework, ISO 42001, and for EU-exposed firms the AI Act — exist precisely to structure these answers. Mapping to them is far cheaper when done during design.
What to do differently: produce model documentation, data lineage and decision logging as build artefacts rather than as a compliance exercise afterwards. Assume from day one that someone will ask the system to explain itself. More on governed architecture →
Ask a stalled project what the baseline was — how long the task took before, how often it was wrong before — and the answer is frequently that nobody recorded it. Without that, the AI system cannot be shown to have helped, which makes it impossible to defend at budget time and impossible to improve deliberately.
The related gap is the evaluation harness: a repeatable test suite measuring output quality against a known-good dataset. Without one, any prompt change or model upgrade is a leap of faith, so teams stop changing anything, and the system slowly falls behind.
What to do differently: record the pre-AI baseline before the pilot begins, and build the evaluation harness as part of the first build rather than a later improvement. It is unglamorous and it is what makes everything after it safe.
Each of these failures is a version of the same thing: the project was structured around producing a demonstration rather than transferring a working capability to the people who will live with it.
That framing also gives a useful test when evaluating any external partner. Ask what you will own at the end. Ask whether the methodology will be taught or held. Ask which of their systems your workflow will depend on after they leave. The answers separate firms that build toward your independence from firms that build toward your continued need for them.
The bottleneck in mid-market AI is not enthusiasm and it is not access to models. It is the distance between a pilot and an owned system — and that distance is made of ownership, data, governance and measurement.
If one or more of these is familiar, the useful first move is usually not another pilot. It is a structured diagnosis of which of the five is actually blocking you, since the remedies are very different and expensive to guess at. That is what our two-week readiness assessment is built to answer.
Start here
Describe where the project stopped and we can usually tell you which of the five it is within a conversation.