Insights · Operating system

Why most AI pilots stall, and what the ones that scale do differently

Four out of five enterprise AI pilots never reach production. The reasons are rarely technical.

Most executives have now run at least one AI pilot. Most of those pilots produced a demo that impressed the room, a slide that made it into the board pack, and very little else. Industry surveys put the share of pilots that reach production at somewhere between one in five and one in ten. The pattern is consistent enough that it deserves a diagnosis rather than a shrug.

The pilot was designed to impress, not to run

A pilot that starts from a prompt and an export of last quarter's data answers one question: can the model do this at all? It almost always can. What it does not answer is whether the system can run every day, inside the tools the team already uses, on data that changes, with the exceptions a real operation produces. Those are the questions that decide whether anything ships, and a demo is structurally unable to answer them.

Each pilot starts from zero

The second pilot in most companies takes as long as the first, because nothing carries over. The data was extracted by hand for one use case. The rules lived in a prompt. The people who knew what "an approved order" meant explained it to the vendor once, verbally, and then the vendor left. The company owns a demo, not an asset.

The pilots that scale look different from the start. They are built on a model of the business that outlives the use case: the objects the company already talks about, the rules that bound what can be done, and the evaluations that say how good a decision is. The first use case takes longer to build that way. The second ships in a fraction of the time, because the model is already there.

Nobody owned the number

Ask who is accountable for the pilot's result and you usually get a function, not a person and a P&L line. Pilots that scale are scoped the other way round. They start from a line an executive already owns, such as days sales outstanding, quote turnaround, or long-tail spend, and they agree a baseline before anything is built. When the number moves, the case for production makes itself. When it does not, the pilot stops, which is also a good outcome.

The system was never allowed to act

Many pilots are read-only by design: they summarise, they recommend, they draft. The return on that is real but small, because a person still has to do the work. Systems that pay back are allowed to act, through a gate a person controls: approve the invoice, place the order, send the quote. Getting there needs shadow mode on live data, an audit trail with sources, and rules the business can read. It is more work than a demo. It is also the only version that changes the economics of the operation.

What to ask before the next pilot

  • Which P&L line will this move, and what is the baseline today?
  • Will it run inside the tools the team already uses, on live data, in our own tenant?
  • What will carry over to the second use case?
  • What is the system allowed to do without a person, and who holds the gate?
  • What happens at month six if the number has not moved?

A pilot that has clear answers to those five questions is not really a pilot. It is the first system of many.