A pilot that starts without a named success metric, a named owner, and a decision date will still be running in a year, and by then it will be production.
One statistic frames the whole page. In a survey of 623 technology decision-makers, 72% agreed their organization expects agentic AI to deliver a return, and 5% said they had well-defined success metrics for it, with another 19% reporting “some.” That is an expectation almost nobody is instrumented to test. It is AWS-sponsored and skews large-enterprise, and it is still the most useful number here, because the gap it describes has to be closed at design time rather than afterward.
The checklist
1. Write the value hypothesis and how it will be measured, before launch. A CIO’s version of this is the cleanest formulation in the corpus: define the hypothesis and its measurement before scaling, not after the demo lands well. The corollary is the decision date, since a pilot with no date never fails and simply persists.
2. Choose the use case on shape, not on enthusiasm. The discriminator worth stealing: a workflow with coordination overhead, rigid sequences, frequent human intervention and room for dynamic adaptation is an agent candidate. A highly standard, repetitive workflow with limited variability (payroll runs, expense approvals, password resets) is a script. Putting an agent where deterministic automation works buys nondeterminism that then has to be governed. Agents fit tasks with clear parameters, repeatable structure and measurable outcomes, not decisions requiring judgment or accountability (that framing is an interviewed expert in a sponsored report, not its survey data).
3. Pilot the most reliable case first, deliberately. Early negative encounters send low-reliability signals and reduce trust; that comes from a synthesis of 197 studies. Whether one failed first pilot poisons a whole firm’s appetite is our extension of that finding rather than a measured effect, but the direction attracts little serious dispute, and the cost of testing it the hard way is high.
4. Let the use case originate with a domain expert. Siemens’ data-first “big bang” approach (connect everything, then find uses) failed; what worked started from a domain expert’s problem, and after the first working case more use cases came from frontline staff unprompted.
5. State the counterfactual in the same currency as the investment. The most transferable move in this corpus: one plant weighed a defect-prediction model against buying another €500,000 x-ray tester. That reframes the pilot as a capital-allocation decision, a probabilistic option against a deterministic one, rather than as an experiment competing against nothing.
6. Stagger the rollout, because that is what makes measurement possible. In the most-cited field study of AI at work, the initial pilot covered roughly 50 workers, about half randomized into treatment; the usable treatment arm was 22 people and the authors had no control-group data. What produced the causal estimate was the staggered rollout that followed, forced by a training bottleneck; a rollout constraint is a research design if a firm treats it as one.
7. Don’t automate a broken process. Make it well defined, repeatable and documented first. Practitioner common sense rather than a finding, and from a source whose overall posture (buy or partner, never build) is the opposite of this wiki’s premise.
8. Agree the kill decision before launch, in writing, with a named owner who can make it.
Measurement traps
Seat counts are not usage. In a field experiment across 66 firms and 7,100-plus workers, more than 90% of people assigned the tool used it at least once in six months, while the mean worker used it in 41% of study weeks, 1.6% used it every week, and 10% never used it at all. Weekly use settled just under 40% of licensed workers after an initial novelty peak. Most instructive for a pilot designer: firm identity predicted usage better than anything about the individual, with a cross-firm range from 6.3% to 75%, so a pilot measures the firm rather than the tool.
Don’t expect a team-saturation effect at pilot scale. The same authors looked and found few substantive differences among workers whose close colleagues also had access, concluding that larger shifts in responsibility need time and institutional effort, not local coordination. Saturating a team may still be right; this is evidence against expecting it to show up in a pilot.
Acceptance, adoption and usage are three different things. Headline adoption statistics measure formal commitment. The widely-quoted figure that most firms use AI in at least one function says nothing about intensity and should never be repurposed as though it did.
Time-saved-per-task dashboards miss scope expansion: people take on work they previously wouldn’t have. That observation comes from researchers employed by the vendor whose product they studied, a conflict worth carrying in the same sentence as the observation.
Refuse the circulating ROI figures. A widely-quoted “2.3× return within 13 months” is a vendor-sponsored InfoBrief, self-reported, with 13 months being time-to-first-ROI rather than payback. A widely-quoted “$3.50 back per $1 invested in agentic AI” is neither the consultancy’s own research nor agentic-specific; its footnote points to a 2023 study of AI investment generally, written before agents.
The 30/60/90 question
The frame is under attack from the evidence. One 2026 benchmarking study of 1,221 technology leaders reports that roughly three-fifths of firms take up to twelve months to get AI workloads into production, and argues explicitly that a 30/60/90 cadence falls well short. Only the three-fifths shape is independently corroborated; the finer banding sits behind a download gate, and the sponsor sells data infrastructure, which makes the study a serious challenge rather than a settled refutation.
Deloitte’s version is more concrete and names the mechanism an anonymous executive there calls pilot fatigue: pilots run small, clean and isolated, while production demands infrastructure, integration, security review, compliance, monitoring and maintenance. Use cases scoped for three months stretch to eighteen once that complexity arrives (an assertion in a narrative, with no distribution behind it). Only 25% of firms had moved more than 40% of experiments into production, and 54% expected to within three to six months, which is intent rather than forecast.
Our position: keep the 30/60/90 for the decision cadence and stop pretending it is the production timeline. Day 30 checks whether the thing works at all, day 60 whether the measurement is real, day 90 forces kill-or-promote. Promotion then enters a gate whose clock is set by integration reality, not by the pilot calendar.
How you’d know it’s working
Pilots end. A healthy program has a record of pilots killed at day 90; read that record for which ones died. Killing three throwaways while the one with an executive sponsor runs eighteen months satisfies the count and none of the intent. A portfolio of perpetual experiments is agent sprawl with better branding.
There is a measured baseline-versus-pilot delta, not a testimonial. Reuse the pilot’s success criteria as the acceptance evals; if they can’t serve as an eval, they weren’t criteria.
Benefits realized track benefits sought. In the survey above, organizations sought productivity (48%), better decisions (41%) and cost savings (39%), and reported achieving them at 36%, 35% and 33%, closer than the pilot-failure literature would suggest. The revealing split is the 10% reporting no benefit at all: 3% among readiness leaders against 20% among followers.
What this doesn’t solve
A successful pilot proves the use case. It leaves both the platform and the governance unproven: pilot-scale controls buckle at production volume, and promotion still goes through the gate.
It also cannot resolve a genuine disagreement about how much process a pilot deserves. One prominent practitioner argues for building fast and dirty with cross-functional teams, releasing into the organization, measuring and repeating, which cuts directly against gating before release. The reconciliation we offer is gate depth by tier, and it is worth being plain that he would still disagree at the low tiers.
There is a second unresolved tension inside the selection rule. “Start with quick wins and clear ROI” is exactly the instinct that one study blames for funding visible, low-payback use cases while the higher-return back-office opportunities stay starved. Both cannot be the selection rule. Our read: use quick wins to build the capability, and choose the second wave on measured return rather than visibility.
Finally, the maturity context. Only about 1% of firms self-rate as mature on AI while 92% plan to increase investment. Whatever the pilot proves, it is being read by an organization that has not done this before.
See also
- Day 3 sequencing — where pilots sit in the program sequence.
- Evals — the acceptance tests a pilot’s success criteria should reuse.
- Promotion gates — what stands between a good pilot and production.
- Agent sprawl — what perpetual pilots turn into.
- Risk-tier assignment — the axis that sets how much process a pilot deserves.
- Build vs buy — the platform decision a pilot should not be asked to settle.