Claim a free live session
← The feed
The stall

Why enterprise AI pilots fail, and what production actually requires

The model was never the constraint. Seven forces hold the gap between a working demo and a department that runs on it — and a diagnostic clears none of them.

Michael HalasCo-founder · A³9 min read
An engineer working through a problem at a screen on the A³ campus

The demo landed. The room agreed it was impressive. The deck circulated to two more committees, someone asked about data residency, someone else asked about cost at scale, and the answers were good enough that nobody objected. Six months later the tool sits next to the workflow instead of inside it, the budget clock is still running, and the question at the next review is no longer "does this work" but "what happened."

If that is recognisable, the useful thing to know is that it is not a technology outcome. It is the default outcome. And almost nothing about fixing it involves choosing a better model.

The failure is not technical

There are two different capabilities in play, and the industry keeps collapsing them into one word. The first is model capability: can the software do the task at an acceptable quality. The second is organisational capability: can this department run the task this way, every day, without the vendor in the room.

Model capability is now rarely the binding constraint. It is testable in an afternoon, and if it fails you find out early and cheaply. Organisational capability fails late, quietly, and expensively — which is why so many programmes look healthy right up to the moment they are cancelled.

A pilot proves the software can do it. Nothing in a pilot proves your organisation can.

The seven forces that hold the gap open

Each of these is diagnosable from where you sit today. Read them as a checklist against your own programme rather than as commentary.

  • Vendor noise, and the cost of choosing wrong. Every category has a dozen credible-looking suppliers and no reliable way to separate the ones who deploy from the ones who demo. The symptom: your shortlist was assembled from search results, analyst grids and warm introductions rather than from anything you watched run.
  • No named workflow. "AI for the finance team" is a mandate, not a workflow. Until a specific process is named, with a person who owns its number, there is nothing for a deployment to attach itself to. The symptom: nobody can state, in one sentence, which task will be done differently on Monday.
  • Talent scarcity. The people who can carry a first deployment across the messy middle are rare, expensive, and usually already committed. The symptom: the pilot depended on one internal enthusiast, and it slowed the moment they were pulled elsewhere.
  • Training that produces certificates. Generic literacy programmes teach a category, not a job. The symptom: the team completed the course and the workflow did not change.
  • Board pressure on a moving clock. The mandate arrives with an expectation of visible progress on a quarterly rhythm, which pushes teams toward announceable activity rather than compounding capability. The symptom: your roadmap is optimised for the next board slide.
  • A governance surface that grows with the rollout. What was fine for twelve users in one team becomes a legal review at two hundred across four. The symptom: governance was consulted after the pilot succeeded rather than before it scaled.
  • Nobody owning the outcome across vendors. The model vendor owns the model, the integrator owns the integration, the platform team owns the platform, and the outcome is owned by a committee. The symptom: when it stalls, every party can honestly say their part worked.

What production actually requires

Strip the seven forces down and four requirements remain. They are unglamorous and they are the whole game.

  • A named workflow with a P&L owner — someone whose number moves if this works.
  • A partner matched to that specific problem, not to your industry in general.
  • The people who will operate it present while it is being built, not trained afterwards.
  • A governance line drawn before scale, at the risk tier the workflow actually sits in.

Why another diagnostic clears none of this

The market's standard answer to a stalled programme is a free assessment. It is genuinely good at one thing: naming the gap. But you already felt the gap — that is why you are reading this. An assessment ends where the difficulty begins, and it ends there by design, because its commercial job is to open a consulting engagement rather than to ship a workflow.

A diagnosis has never shipped anything. It tells you that you are unfit; it does not do the training.

What proof looks like instead

The alternative is narrower and harder to fake: the capability running live, on a problem shaped like yours, in front of the team that owns the number. Not a curated demo on clean data with a motivated operator — a working session where the failure modes are visible and the questions come from people who will have to live with the result.

That format does something an assessment structurally cannot. It lets the people who will run the tool judge it before the organisation commits, and it produces a shared memory that survives the next reorganisation.

Your next move

It is not commissioning another audit. It is naming the problem precisely enough that someone can demonstrate against it — the workflow, the number it affects, and who owns that number. That sentence is the entire qualification. Everything downstream, including which partner you should be talking to, follows from getting it written down.

Reading helps. Seeing the capability tested is better.

When you're ready to move from ideas to evidence, bring us the problem and we'll build the right session around it. The first one is free.