Why AI pilots stall — an operational post-mortem
Six failure patterns we see repeatedly in stalled AI projects. None of them are about the model. All of them are visible in the first fortnight if you know what to look for.
We get called in after the fact more often than before it. A business ran a pilot, it generated a fortnight of enthusiasm, and then it quietly stopped being mentioned. Nobody can say precisely when it died or why.
Across those post-mortems the same patterns recur. Almost none of them are about the technology. Here they are, roughly in order of how often we see them.
1. Nobody owned it after launch
By some margin the most common cause.
The pilot has a sponsor — usually someone senior and enthusiastic — and a delivery team, often external. It goes live. The delivery team leaves. The sponsor moves on to the next thing. And the tool sits there, gradually diverging from reality as prices change, staff change and processes change, until using it becomes actively worse than not using it.
A sponsor authorises. An operator maintains. These are different people with different incentives, and a pilot with only the first is on a timer from the day it launches.
What to look for in week one. Ask who will make the change when a price goes up. If the answer is a vendor, or a name nobody can produce, you have found the problem before it costs you anything.
2. There was no baseline
The pilot cannot succeed because success was never defined in measurable terms.
We see this constantly. The project was justified with “improve customer experience” or “increase efficiency” and nobody wrote down what those were before the pilot started. Six months later somebody asks whether it worked, and the honest answer is that nobody can tell.
Worse, without a baseline the pilot is vulnerable to whoever has the strongest opinion. One memorable complaint from a senior person can kill something that was quietly working, because there is no number to defend it with.
What to look for. Ask what the current number is. Calls answered, hours spent, tickets resolved, conversion rate — anything, as long as it was measured before the change. If nobody measured it, stop and measure it. Two weeks of baseline is worth more than two months of pilot.
3. It was never connected to anything
The demo worked because it lived in its own world. Then it went live and had to interact with the booking system, the CRM, the inventory, and the way the operations team actually work.
Integration is where the effort is, and it is systematically underestimated because it is the least interesting part to talk about. A pilot that cannot write to the system of record is not a pilot. It is a very expensive way of generating work for whoever has to re-key the output.
What to look for. Ask what happens to the output. If a human copies it somewhere, the pilot has not been tested at all — you have tested the model and skipped the operation.
4. It solved a problem nobody had
Someone saw a capability and looked for somewhere to apply it. This produces technically successful pilots that change nothing, because they optimised a step that was not a constraint.
Automating a task that takes forty minutes a week is a fine engineering exercise and a waste of capital. The thing worth automating is the one your team complains about, not the one that demos well.
What to look for. Ask which person’s job gets easier and by how much. If nobody can name them, or the answer is “everyone, a bit”, it is a demo.
5. It was launched at full scope
A pilot that goes live across every channel, for every customer, on day one has no safe way to fail. The first bad interaction becomes an organisational event rather than a tuning note, and the political cost of the second one is usually fatal.
The pilots that survive start somewhere the downside is small. After hours, where the alternative is voicemail. One product line. One branch. They accumulate evidence before they accumulate exposure.
What to look for. Ask what happens the first time it gets something wrong in front of a customer. If the answer is “we would have to switch it off”, the scope is too broad.
6. The people using it were told, not asked
The pilot was designed with management and delivered to staff. The staff were not consulted about the process being changed, were not asked what actually slows them down, and quite reasonably concluded this was something being done to them.
They then do not use it, or use it minimally, and the pilot fails on adoption while everyone looks at the technology.
What to look for. Ask whether the people who will use it were in the room during design. Not in a briefing after design — in the room during it.
The uncomfortable pattern
Read those six back and notice what none of them are. None of them are about model quality, prompt engineering, hallucination rates, or which vendor was chosen. The technology is, in almost every post-mortem we have run, the part that worked.
That is genuinely good news, because it means the failure modes are ordinary project-management and operations failures, and those are well understood. It also means the fix is not a better model. It is the boring discipline of ownership, measurement, integration and scope — applied before the build rather than after it.
What we do differently
This is why our roadmap engagement spends more time on the operation than on the technology, and why our builds start with the narrowest viable scope rather than the most impressive one.
It is also why we insist on measures agreed at kick-off. Not because it makes us look rigorous, but because a project without them cannot be defended six months later, and undefendable projects get switched off regardless of whether they were working.
If you have a stalled pilot and want an outside read on why, we will do that as part of a readiness engagement — and if the answer is that it should be switched off rather than rescued, we will tell you that too.