Adoption · 8 min read · Aug 15, 2026
Most agent pilots die before production.
There are two numbers circulating about enterprise AI agents in 2026. One says more than half of organisations are deploying them. The other says 88% of pilots never reach production. Both are measured, both are true, and the gap between them is the whole story.
The adoption numbers disagree, and both are right
KPMG's Q1 2026 AI Pulse survey puts 54% of organisations actively deploying AI agents across core operations, up from 11% two years earlier. Separate work from S&P Global Market Intelligence and McKinsey puts the figure at 31% with at least one agent genuinely in production. The two are not in conflict — they measure different things. One counts intent and activity, the other counts systems that survived contact with real work.
The sector spread inside that second number is wider than the headline: banking and insurance lead at around 47%, while healthcare sits near 18% and government near 14%. Regulation explains most of the gap, and it explains it in a way that is worth taking seriously rather than treating as backwardness.
There is also a third number that catches operators by surprise. Gartner predicts 40% of enterprise applications will feature task-specific AI agents by the end of 2026, up from less than 5% in 2025. Which means a meaningful share of groups will soon run agents they never explicitly chose — they arrive in a vendor update. If you have not audited that, you have an unmanaged surface, not an absence of one.
The three blockers, and which one actually matters
Where pilots stall, leaders cite evaluation gaps (64%), governance friction (57%) and model reliability (51%). Broader survey work is bleaker still: only around 25% of AI initiatives deliver the ROI expected of them, and only 16% reach enterprise-wide scale.
Reliability gets the attention because it is the most visible — the model did something odd, everyone saw it. Governance gets the budget because it produces documents. But evaluation is the one that decides outcomes, and it is the least glamorous of the three.
An evaluation gap means nobody agreed, in advance, how to tell a good output from a bad one. Without that, a pilot cannot be passed or failed — it can only be argued about, and arguments do not get deployed.
Why evaluation is the real gate
Ontilus has written elsewhere that a consolidated number is worthless until it reconciles against the outlet's own close report. Agent work has exactly the same shape. If an agent codes an invoice, there has to be a set of invoices where the right answer is already known, and a threshold agreed before anyone runs it.
This is unglamorous and it is the difference between a pilot that concludes and one that drifts. A pilot without a pass mark has no natural end: it produces demos, the demos are impressive, somebody senior asks whether it is trustworthy, and there is no answer that isn't an opinion. Six months later it is quietly deprioritised. That is what most of the 88% look like from the inside — not a dramatic failure, just a slow absence of evidence.
What the ones that survive have in common
Across the deployments that graduate, four things recur:
- A bounded process, not a capability. "Code supplier invoices to the right GL account" graduates. "Use AI in finance" does not. The scope has to be small enough that correctness is decidable.
- A checkable output. There is a ground truth to compare against — last quarter's coded invoices, the close report, the approved roster. If nothing can be checked, nothing can be trusted.
- A named owner with authority to stop it. Governance friction is usually not excessive process; it is the absence of anyone empowered to decide, so the decision escalates and stalls.
- A human step that is real. Review that nobody performs is worse than no review, because it manufactures the appearance of control. Either the human genuinely checks, or the process is autonomous and monitored as such.
What to demand before you start
Three questions, asked before any build, filter most of the failure out. What is the ground truth we will measure against, and does it already exist? What accuracy would make this worth deploying, stated as a number, agreed by the person who owns the process? And what happens on the day it is wrong — who notices, and how?
A vendor who cannot answer those has not scoped the work. Neither have you, and the pilot will be one of the 88%.
Sources
- KPMG AI Pulse Survey, Q1 2026 — 54% actively deploying agents.
- S&P Global Market Intelligence and McKinsey, 2026 — 31% with at least one agent in production; sector breakdown.
- Gartner press release, Aug 2025 — 40% of enterprise applications will feature task-specific AI agents by 2026, up from under 5% in 2025.
- 2026 enterprise agent pilot surveys — 88% of pilots fail to reach production; blocker percentages.
Dealing with this in your own group?
We answer scoping questions before there's a contract in sight — including the ones about cost and data handling.