Adoption · 8 min read · Aug 15, 2026

Most agent pilots die before production.

There are two numbers circulating about enterprise AI agents in 2026. One says more than half of organisations are deploying them. The other says 88% of pilots never reach production. Both are measured, both are true, and the gap between them is the whole story.

The adoption numbers disagree, and both are right

KPMG's Q1 2026 AI Pulse survey puts 54% of organisations actively deploying AI agents across core operations, up from 11% two years earlier. Separate work from S&P Global Market Intelligence and McKinsey puts the figure at 31% with at least one agent genuinely in production. The two are not in conflict — they measure different things. One counts intent and activity, the other counts systems that survived contact with real work.

The sector spread inside that second number is wider than the headline: banking and insurance lead at around 47%, while healthcare sits near 18% and government near 14%. Regulation explains most of the gap, and it explains it in a way that is worth taking seriously rather than treating as backwardness.

There is also a third number that catches operators by surprise. Gartner predicts 40% of enterprise applications will feature task-specific AI agents by the end of 2026, up from less than 5% in 2025. Which means a meaningful share of groups will soon run agents they never explicitly chose — they arrive in a vendor update. If you have not audited that, you have an unmanaged surface, not an absence of one.

Of every 100 enterprise AI agent pilots, 12 reach production. The most-cited blockers are evaluation gaps at 64 percent, governance friction at 57 percent, and model reliability at 51 percent.
Reported blockers among leaders whose agent pilots stalled. Evaluation leads, and it is the one most proposals leave out.

The three blockers, and which one actually matters

Where pilots stall, leaders cite evaluation gaps (64%), governance friction (57%) and model reliability (51%). Broader survey work is bleaker still: only around 25% of AI initiatives deliver the ROI expected of them, and only 16% reach enterprise-wide scale.

Reliability gets the attention because it is the most visible — the model did something odd, everyone saw it. Governance gets the budget because it produces documents. But evaluation is the one that decides outcomes, and it is the least glamorous of the three.

An evaluation gap means nobody agreed, in advance, how to tell a good output from a bad one. Without that, a pilot cannot be passed or failed — it can only be argued about, and arguments do not get deployed.

Why evaluation is the real gate

Ontilus has written elsewhere that a consolidated number is worthless until it reconciles against the outlet's own close report. Agent work has exactly the same shape. If an agent codes an invoice, there has to be a set of invoices where the right answer is already known, and a threshold agreed before anyone runs it.

This is unglamorous and it is the difference between a pilot that concludes and one that drifts. A pilot without a pass mark has no natural end: it produces demos, the demos are impressive, somebody senior asks whether it is trustworthy, and there is no answer that isn't an opinion. Six months later it is quietly deprioritised. That is what most of the 88% look like from the inside — not a dramatic failure, just a slow absence of evidence.

What the ones that survive have in common

Across the deployments that graduate, four things recur:

  • A bounded process, not a capability. "Code supplier invoices to the right GL account" graduates. "Use AI in finance" does not. The scope has to be small enough that correctness is decidable.
  • A checkable output. There is a ground truth to compare against — last quarter's coded invoices, the close report, the approved roster. If nothing can be checked, nothing can be trusted.
  • A named owner with authority to stop it. Governance friction is usually not excessive process; it is the absence of anyone empowered to decide, so the decision escalates and stalls.
  • A human step that is real. Review that nobody performs is worse than no review, because it manufactures the appearance of control. Either the human genuinely checks, or the process is autonomous and monitored as such.

What to demand before you start

Three questions, asked before any build, filter most of the failure out. What is the ground truth we will measure against, and does it already exist? What accuracy would make this worth deploying, stated as a number, agreed by the person who owns the process? And what happens on the day it is wrong — who notices, and how?

A vendor who cannot answer those has not scoped the work. Neither have you, and the pilot will be one of the 88%.

Sources

  • KPMG AI Pulse Survey, Q1 2026 — 54% actively deploying agents.
  • S&P Global Market Intelligence and McKinsey, 2026 — 31% with at least one agent in production; sector breakdown.
  • Gartner press release, Aug 2025 — 40% of enterprise applications will feature task-specific AI agents by 2026, up from under 5% in 2025.
  • 2026 enterprise agent pilot surveys — 88% of pilots fail to reach production; blocker percentages.

Dealing with this in your own group?

We answer scoping questions before there's a contract in sight — including the ones about cost and data handling.

Questions

Short answers,
in full.

The questions this article gets asked most, answered so each one stands on its own.

Talk to us
What percentage of AI agent pilots reach production?

Survey data from 2026 puts it at roughly 12% — 88% of enterprise agent pilots do not graduate to production. Adoption figures that look much healthier, such as the 54% of organisations reported to be actively deploying agents, are measuring activity rather than systems running in production. A narrower measure of at least one agent genuinely in production sits nearer 31%.

What is the most common reason agent pilots fail?

Evaluation gaps, cited by 64% of leaders — no agreed method for deciding whether an output is good. Governance friction follows at 57% and model reliability at 51%. Evaluation is the most consequential because without a pass mark agreed in advance, a pilot cannot conclude; it produces demonstrations rather than evidence and is eventually deprioritised.

How do we know if an agent is accurate enough to deploy?

Decide the threshold before building, against data where the correct answer is already known — previously coded invoices, signed-off close reports, approved rosters. The process owner states the accuracy that would make deployment worthwhile, and the pilot is measured against it. An accuracy target set after seeing results is not a test.

Are we already running AI agents without knowing?

Increasingly likely. Gartner predicts 40% of enterprise applications will feature task-specific AI agents by the end of 2026, up from less than 5% in 2025, and expects most enterprise apps to already carry AI assistants — the precursor to agents. Agents commonly arrive through a vendor update rather than a procurement decision, so auditing what your existing software already does autonomously is usually more urgent than adding anything new.