Builders under-tier their own agents, so tier assignment must be a process with a scorer who is not the owner, a written rubric, and a re-score trigger on every capability change.
This page is the assignment process. The tier model itself, meaning what the tiers are and why autonomy × authority × criticality is the scoring axis, is canonical in risk tiers and trust zones.
The checklist
1. Score at registration, automatically, and publish the rubric. Arcadis routes on a questionnaire filled in when a builder registers the project: answers about funding, data sources and tools sort the build into a band, and the bands are published to builders on purpose. Microsoft’s own account of the same program describes the mechanism concretely: solutions intended for more than a hundred users go to the centre of excellence for deeper assessment, smaller ones get an environment provisioned automatically. That is the cheapest implementable version of this page. (The book adds a green/yellow/red/black colour scheme and a headcount unsupported by Microsoft’s write-up; use the mechanism, not the trimmings.)
2. Score the combination, not one axis. WEF’s framework is explicit that impact emerges from the interaction of role, autonomy, authority, predictability and operational context: high autonomy in a simple context carries different risk from moderate autonomy in a complex one. A rubric that scores only “does it touch client data” will systematically misprice the agent with wide latitude over unremarkable data.
3. Use discriminators that survive contact with a real build. Four earn their place:
Declarative or evaluative. Microsoft’s red-team taxonomy separates an agent following a constrained, user-defined path from one given a goal and latitude. An evaluative agent has discretion the reviewer cannot enumerate in advance, which is a sharper risk signal than data sensitivity.
Is failure auditable. Practitioners in a 16-interview field study named this as the deployability test, alongside whether a guardrail or manual gate can be inserted at the quality-critical stage.
Reversible or not. Reversible, low-impact actions tolerate retrospective review; irreversible or financial ones need approval before execution. That formulation comes from a legal op-ed rather than a study, so take it as a usable line, not as authority.
Recovery cost. MATRA supplies the most directly reusable rubric in this corpus. It asks, per asset and per confidentiality/integrity/availability dimension, what the business impact is if that property is compromised, graded Low, Moderate or High by recovery cost: limited and quickly mitigated, versus prolonged or impossible, then crossed with likelihood in a three-by-three matrix to yield a 1–9 rating.
4. Tier on inherent risk; re-tier on residual. WEF’s distinction is worth implementing literally. Inherent risk is likelihood × impact before mitigation. Residual risk reflects what the controls actually demonstrated, with evaluation evidence as the bridge: pass rate on the tier’s eval suite, behaviour under adversarial input, observed error rate in production. Installing a guardrail earns an agent no tier reduction. Demonstrating through evals that the guardrail works does.
5. Pre-commit the re-score triggers. New tool, new data source, new egress channel, new audience: each forces a re-score before the change ships. The structural idea worth stealing from the AI-control literature is the if-then tripwire: decide in advance which controls turn on at which threshold, so that nobody renegotiates the rubric during an incident. Take the structure only; that paper is about misalignment as a threat model, excludes human misuse, and its upper levels are well ahead of practice.
6. Separate scorer from owner, and give disputes somewhere to go. The cheap version is a three-question rubric applied by the platform lead. The thorough version rotates scorers, samples completed scores for audit, and records the rationale. Without an appeal path, a contested score becomes an argument for not registering next time.
7. Write the carve-outs down. Some actions no tier authorizes. A healthcare operator running three-tier “directed autonomy” in production still states plainly that an agent will never produce a diagnosis alone. Truist’s generative-AI lead makes the finance version of the point: he doubts any genuinely high-risk use case in financial services should run fully human-out-of-the-loop. Name the local equivalents (trade execution, client communication, wire initiation) before someone asks whether tier 3 covers them.
8. Tier the use case before tiering the build. Risk appetite can be implemented by choosing what gets built at all. Dentsu’s citizen program stayed lightly governed largely because the portfolio was deliberately low-risk work. Citi expects supervisors to ask where a firm chose to deploy agents and why, and to read the answer as a risk-appetite statement. The rubric is what answers that question in writing.
One regulated firm, an unnamed nuclear utility in Davenport and Barkin’s fieldwork, spent nine months designing a three-tier ladder before launching citizen development at all: tier 1 personal productivity with the loosest controls, tier 2 triggered by multi-user reach or production-system contact, tier 3 requiring SOC standards assessed by independent auditors. Unverifiable outside the book, and the structure matches what Microsoft’s maturity guidance prescribes independently; cite it as one firm’s account rather than as prevalence.
How you’d know it’s working
Tier-change events per quarter, above zero. A portfolio where nothing ever moves tiers indicates that re-scoring has stopped, rather than that agents are stable.
Audit-sample disagreement rate. Re-score a sample blind and compare. Systematic downward disagreement between owner-adjacent and independent scorers is the effect this page exists to counter, and it is measurable.
Utilization of the middle tier. The middle rung fails first. Arcadis’s assisted tier was the one that mostly went unused, because the fusion-team capacity to staff it never existed. Where a firm’s middle tier is empty, the question is whether it is unnecessary or unstaffed, and those have opposite fixes.
Whether the autonomy ceiling tracks domain risk rather than model capability. In the field study, the single firm operating at the highest autonomy level got there because it worked in a low-risk analytic domain with no formal qualification requirements, rather than because it was more technically advanced. Where the highest-autonomy agents in a portfolio are the most sophisticated rather than the least consequential, the rubric is being scored on the wrong thing.
What this doesn’t solve
A tier label does nothing by itself. It pays off only when gates, controls and review cadence actually differ by tier; tiers plus uniform treatment is paperwork. Microsoft’s own guidance names the failure directly, though as a vendor’s assertion rather than a measurement: treating all agents alike either over-restricts low-risk ones, which pushes builders into the shadows, or under-governs mission-critical ones.
The rubric prices risk without controlling what work arrives. Agents carry a high fixed cost (specifying the objective, reviewing the output) and a low marginal cost per step, so they win decisively on long, multi-tool tasks. The traffic self-routes complex work toward agents regardless of how it is scored. Expect the portfolio to drift upward in tier and staff for that. (The economic framing comes from Perplexity-affiliated researchers studying Perplexity telemetry, disclosed; follow their own instruction to cite direction and not magnitude.)
Assigning “human in the loop” as a tier outcome is weaker than it reads. The NBER work on human-AI collaboration finds humans under-respond to AI predictions and reduce their own effort when shown confident ones, and that the incremental value of assisting a human with a prediction is small next to automating the decision outright. A tier whose only control is a human reviewer has bought less than the rubric claims; see oversight decay.
And tiering governs the registered set. Nothing here reaches an agent nobody declared; see shadow agents.
See also
- Risk tiers and trust zones — the model this process applies.
- Three-zone architecture — where tier boundaries become architecture rather than labels.
- Promotion gates — the gate evidence keyed to each tier.
- Agent inventory and registry — where the tier score is recorded and re-scored.
- Templates and checklists — the tiering rubric as a printable artifact.
- Oversight decay — why “human in the loop” is a weak tier outcome.