July 31, 2026
Most AI pilots die at the switch, not in the sandbox
One in three small businesses using AI is stuck in testing. The reason is not skill or budget. It is that the tool only has two settings.
Pax8 fielded a survey of 402 US small business leaders this quarter. Sixty-one percent are actively using AI. Another 29 percent are still experimenting. The number worth sitting with is this one: nearly one in three of the businesses already using AI are stuck in experimentation, unable to move anything into real operation.
That is not a small business problem. On July 28, Cognizant launched a dedicated EMEA unit built around closing the same gap, citing IDC research that 88 percent of AI agent proofs of concept never reach broad production. For every 33 pilots a company starts, four go live.
Asked why they are stuck, the owners in the Pax8 survey point to lack of internal expertise, unclear ROI, and security worries. Those answers are honest and they are also symptoms. The disease is in the shape of the tool.
Gartner said the quiet part in May
Gartner’s Shiva Varma put it plainly when the firm published its agent governance research: companies treat agent oversight as either locked down or fully trusted, “and that is the root cause of failure.” The firm expects governance gaps to push 40 percent of enterprises into demoting or decommissioning their autonomous agents by 2027, and the gaps typically get discovered after something goes wrong in production.
Two settings. Off, or answerable for everything.
Look at what that does to a pilot. A locked-down trial runs in a sandbox on sample questions, watched by the person who set it up. It produces a demo. What it cannot produce is evidence about the only question the owner actually has, which is whether this thing holds up on a bad Tuesday, on the messy real question from the customer who is already angry, at 11pm when nobody is watching.
So the pilot runs. And runs. It keeps not answering the question, because it was never built to answer it. Then someone schedules the go-live meeting and the owner is standing at a switch with a demo on one side and full public liability on the other. Refusing to flip it is not timidity. Given the evidence available, it is the correct call.
The pilot worked perfectly. It just answered a question nobody was asking, and left the only one that mattered untouched.
The ladder helps. It is still too coarse.
The fix everyone is converging on is graduated autonomy. Gartner recommends four tiers, from read-only observation up through act-with-approval to full independence. Singapore’s framework uses five. Both share the good instinct: promote the agent when its own performance logs earn the promotion, not when the calendar says the pilot is over.
Here is where I think the frameworks stop one notch short. They govern the agent.
An agent is not one thing that is uniformly good. The same system that handles “where is my order” all day without a stumble is on far thinner ice with “the custom piece arrived damaged, I am past the return window, and I want this fixed today.” Same agent, nowhere near the same risk. Put a single autonomy setting across that spread and you get to choose which mistake to make: hold back the work it has clearly earned, or release the work it has not.
Both choices are bad, and the second one is how a company ends up in Gartner’s 40 percent.
The unit of trust should not be the agent. It should be the topic.
What that looks like in practice
At IMCeleste, Celeste starts on a business’s own past support history and drafts real answers to real incoming questions while sending none of them. That is a pilot with production traffic and zero exposure. The owner reads what she would have said, on their actual customers, in their actual voice, before any of it counts.
Then it goes live one topic at a time, and the owner decides which. Order status might graduate first, if the record earns it. Refunds can wait until they do. Each graduation is a small reversible decision backed by that topic’s own track record, not a single irreversible bet on the whole agent. And underneath, the model never gets to grade its own homework: Celeste proposes an action, hard code decides whether it is permitted. You can read how the whole loop works or what she actually does day to day.
None of that is a clever feature. It falls out of taking the failure seriously. If the thing standing between a business and a working AI agent is a cliff, the useful work is not coaching people to jump. It is building stairs.
So stop treating the pilot as a phase to escape. Done right it is the bottom step, permanent, because every new topic starts there. The switch you were afraid to flip is gone, and trust accrues one topic at a time instead of arriving all at once.