Flutterfrog
← Journal

5 August 2026 · 4 min read

Why most corporate AI pilots produce nothing

Most companies have run an AI pilot by now. Very few changed how any work actually gets done.

MIT's Media Lab studied this across roughly 150 leadership interviews, 350 employee surveys and 300 public deployments, and published it as The GenAI Divide: State of AI in Business 2025. Around 95% of organisations reported no measurable effect on profit and loss. Separately, IDC reports that roughly 88% of AI agent proofs-of-concept never reach broad production.

Those numbers get quoted a lot. The interesting part almost never does.

It was not the models

The failures were ordinary and organisational. Poor fit into how work actually flows. Tools handed to people who were never given a reason to change what they did on Monday. And above all: no agreed definition of what a good output looked like before anyone started building.

If nobody wrote down what correct means, nothing can be checked. If nothing can be checked, nobody trusts the output enough to send it. So it stays a demo.

That last step is the one that kills projects. Not a dramatic failure — a slow one. The output is plausible, nobody is quite willing to put their name on it, so a senior person reviews every item by hand. Which is exactly the work you were trying to remove.

What a standard actually looks like

Take a quotation leaving your office. Ask what makes it correct, and you get answers like these:

  • Every line priced from the current price list, not a superseded one
  • Tax treatment matching the customer's registration and delivery state
  • Payment terms as agreed with this customer, not the default
  • The clauses legal mandates, present and in the approved wording
  • Specification matching the enquiry, with deviations listed rather than buried

That is twelve to twenty checks for most businesses. They already exist — they live in the head of whoever does the final read before a quotation goes out. They have simply never been written down, which means no system can enforce them and no new hire can learn them quickly.

Why this is the unglamorous part

Writing the standard down is not an AI problem. It is a conversation with your senior people, and it is genuinely tedious. It is also the step that decides whether anything downstream works, which is why we do it before we touch a model rather than after.

Once it exists, it stops being a document and becomes a check the system runs on every item before a human sees it. Anything failing goes back and is corrected. What reaches a desk has already passed.

One honest caveat: this only works where a right answer exists. Quotations, tenders, submittals, compliance filings — yes. Brand and design work — no. You cannot write a check for taste, and anyone who tells you otherwise is selling something.

That is the whole argument. Predictable does not mean the model is clever. It means the same class of work comes out at the same quality every time, and you can see exactly which check passed and which did not.

Have a system worth building?

Start a project