All articles
AI

Most AI agent pilots never become systems: what separates the ones that do

Nearly every company of any size is now running AI agents somewhere. Far fewer can name what came back. The gap is not model quality, it is that a pilot is a demo and production is a process. Here is the difference, and the sixty-day path across it.

RNM Admin30 September 20268 min read
Most AI agent pilots never become systems: what separates the ones that do

There is a particular meeting happening in a lot of companies this quarter. Someone presents the AI programme. The slide says eleven pilots. The next question, from whoever controls the budget, is which of them are in production and what they returned. The room goes quiet.

This is not a failure of the technology and it is not a failure of the people who ran the pilots. It is a structural problem: a pilot and a production system are different objects, and almost nobody was asked to build the second one.

We deploy this work inside client operations, so we see both the pilots that die and the handful that turn into something the business would notice losing. What follows is what actually distinguishes them.

The short version

  • The pilot succeeded and then stopped, because succeeding was the entire scope. A demo that impresses a steering group has no owner, no fallback, no logging and no budget line. None of those are optional in production.
  • Agent projects fail on process readiness, not model capability. If a human cannot write down how the task is done today, an agent cannot do it tomorrow.
  • The reported adoption numbers and the reported return numbers do not match, and the gap is almost entirely the unglamorous work between the two.
  • Four preconditions predict success: a written process, a measurable outcome, a tolerable failure mode, and a named human owner. Three out of four is not enough.
  • If you do one thing, baseline the metric before you build anything. Most programmes cannot prove value because nobody recorded the before.

Why do most AI agent pilots stall?

Four reasons, in the order we encounter them.

The pilot was scoped to impress, not to run. It was chosen because it demos well: a summariser, a chatbot over the handbook, a draft generator. Demoable and valuable overlap far less than anyone expects. The tasks that are actually expensive in your business are usually tedious, high volume, and visually boring.

Nobody owned it after the demo. The pilot lived with an innovation team, a consultant, or one enthusiastic engineer. Production work needs an owner in the line organisation who is accountable for the output, not for the project.

There was no baseline. Handling time, error rate, cost per transaction, backlog age: if you did not measure these before, you cannot claim an improvement after, and finance will not accept a vibe. This is the same failure mode we described in the KPIs that actually predict business health, just wearing a newer hat.

The underlying process was never written down. This is the big one. An agent needs the rules, the exceptions, the escalation path and the definition of done. In most companies those live in the heads of three people, two of whom disagree. Automating an undocumented process does not produce automation, it produces a faster version of the confusion.

Which processes are actually ready for an agent?

Score any candidate process against four columns before you build anything. This table is the whole triage.

ReadyNot ready yet
DocumentationThe steps, exceptions and escalations are written down and currentIt lives in someone's head, or the doc is two years stale
Outcome measureThere is a number today: handling time, error rate, cost per unit, backlog ageSuccess would be described as "it feels better"
Failure modeA wrong output is caught by a downstream check, or is cheap to reverseA wrong output goes straight to a customer, a regulator or a payment
OwnerA named person in the line organisation whose results it changesA project team, a committee, or "IT"

If a process fails the documentation column, the correct first project is not an agent. It is writing the process down. That is unglamorous and it is the actual bottleneck, and it is why our five-day SOP playbook keeps turning out to be the prerequisite for AI work we were not hired to do.

If a process fails the failure-mode column, it can still be a good candidate, but it needs a human in the approval path, which changes the economics. Be honest about that at the start rather than discovering it at go-live.

What production requires that a pilot does not

Here is the list that gets skipped. Every item is cheap to add at design time and expensive to retrofit.

  • A named owner. One person, in the business, not the project.
  • A fallback path. What happens when the agent is down, rate limited, or refuses? If the answer is "the work stops", you have built a single point of failure into a process that previously had none.
  • Logging you can audit. Every input, output and decision, retained long enough to answer a complaint. You will eventually be asked why the system did something in March. Being unable to answer is a serious problem, and in some sectors a reportable one.
  • A cost ceiling and an alert. Agent spend scales with usage, which means a bug scales your bill. Cap it at the account level and alert at a threshold you set deliberately.
  • A human review sample. Not every output, but a fixed random percentage, reviewed by someone competent, forever. Quality drifts silently otherwise.
  • A change process. Prompts, tools and models change. If a change can reach production without review, you do not have a system, you have a shared document that happens to talk to customers.
  • A security review of what it can touch. An agent with tool access is a user with credentials. Scope it like one. The NIST AI Risk Management Framework is the most usable free starting point, and CISA publishes practical guidance for the smaller end of the market.

A sixty-day plan that actually ends in production

Days 1 to 10: pick one process and baseline it. One, not five. Record the current number for a full cycle. Do not build anything this fortnight. This is the step everyone skips and the reason most programmes cannot prove value later.

Days 11 to 20: write the process down. Steps, exceptions, escalation, definition of done. Have the person who does the work check it. Expect to find two undocumented exception paths and one rule nobody agrees on. Resolving that disagreement is real value even if you stop here.

Days 21 to 35: build the narrow version. The smallest slice that touches a real case, with a human approving every output. Not a demo environment. Real work, reviewed before it lands.

Days 36 to 50: shadow mode. The agent handles the full flow while the human still does the work. You compare. This is where you find the failure modes that a pilot never surfaces, because pilots are run on the happy path by the person who built them.

Days 51 to 60: hand over. The owner takes it, with the logging, the cost cap, the review sample and the fallback path in place. Then measure against the day-one baseline.

If that reads slow, compare it to the timeline of the eleven pilots that produced nothing.

What to measure, and what to stop measuring

Stop reporting: number of pilots, number of employees with a licence, prompts run, hours "saved" by a model that estimated its own savings.

Start reporting: cost per unit of work, error and rework rate, cycle time from request to done, escaped defects, and the human review pass rate. These include the expensive half of the work, which is why they are harder to move and worth reporting.

One more, which is unfashionable and important: staff time spent correcting the system. If it rises faster than the time the system saves, the programme is net negative and everyone can feel it before anyone will say it.

What this means for how you staff

The scarce role is not "prompt engineer". It is someone who can sit with an operations team, extract how the work is really done, write it down accurately, and then hold the line on measurement. That is a business analyst crossed with an operations lead, and it has been undervalued for a decade.

We made the wider version of this argument in AI agents and middle management, and the delivery-side version in AI coding agents in 2026.

Frequently asked questions

How long should an AI agent pilot run before we decide?

Long enough to see a full business cycle including month end, and no longer. Most pilots that run past a quarter are not gathering evidence, they are avoiding the decision. Set the success threshold in writing before you start, then hold to it.

Should we build agents in-house or buy a vendor product?

The same test as any other build-versus-buy question: buy where the process is standard across the industry, build where the process is the thing that makes you competitive. We set out the full framework in build versus buy.

Do small businesses get anything out of this, or is it an enterprise story?

Small businesses get the largest relative gain, because agents cover breadth a small team never had. They also carry more risk, because there is often no second pair of eyes on an output. If that is you, the review sample matters more, not less.

What is the single most common mistake?

Automating a process nobody had written down. It is not close.

Where to go next

If you want a candidate process assessed honestly, including the answer that it is not ready, that is the first half of our business operations consulting work, and anything that needs building sits with custom software development. If you would rather start with the conversation, talk to us.

Work with us on this

Run-Stage Operations

Most of what we are describing here is an operations problem before it is anything else, and that is the work we do most.

Ready when you are

Let's build the next chapter of your business: together.

Tell us where you are and where you want to go. We'll come prepared.