“A senior person will look at it” is not a reviewer pool; a reviewer pool is a named bench with defined skills, rotation, and capacity limits, sized before the gate that depends on it opens.

The writing function got cheap. The review function did not, and it is now the binding constraint. Mollick’s account of building a working artifact without writing a line of code turns on four words: only because I knew enough to check the results. Elsewhere he describes a generated training course as surface-level impressive and free of obvious errors, while seeing instantly that it was too text-heavy and had no knowledge checks. Only someone with domain expertise sees that. Plausible-looking output defeats lay review, so a bench assembled for availability rather than expertise is a queue, not a control. (One expert, one domain, March 2025; treat it as illustration.)

The load is documented, if only qualitatively. An eight-month ethnography of a roughly 200-person technology company (on-site observation plus 40 interviews) describes what its authors call task expansion: engineers absorbing the work of reviewing, correcting and coaching colleagues’ vibe-coded output, so that offloading creation to agents grew rather than shrank their queue. That is one qualitative study; it establishes direction and mechanism, not effect size, and nothing in it sizes a pool. Any staffing arithmetic here is our inference. A longitudinal study of 95 professional engineers names the same shift from the other side, as supervisory engineering work: directing and evaluating and correcting AI output, a new category, with 82% reporting less time spent writing code.

The checklist

1. Name the bench before the gate opens. A gate whose reviewers get identified at submission time will be staffed by whoever is least able to say no.

2. Select on demonstrated verification behaviour, and match domain to artifact. Seniority is a poor proxy, and so is technical review skill when the flaw is substantive rather than syntactic. Where agents let people credibly cross occupational lines, decide in advance who has the expertise to judge the result. (The study recommending that is by Perplexity employees studying Perplexity’s own product, disclosed, and it found financial services had the smallest cross-occupation expansion of any sector, so apply the lane-crossing frame to this audience lightly.)

3. Make the pool cross-functional where risk types diverge. Vanguard’s chief architect for AI/ML describes leaders across business, IT and legal jointly assessing and approving risk levels before anything reaches production, precisely because reputational risk and system-availability risk look different from each seat. Two caveats: it comes from an AWS-sponsored report skewed to large enterprises, and it is a named-executive interview rather than audited process documentation; it is the best asset-manager example available, and still one example.

4. Route intake review to a professional developer. The pre-agent evidence is clear about why: a professional reviewing the proposed use case keeps IT aware of the app (their answer to shadow IT) and stops citizen developers overreaching their technical competence. That is a 2024 low-code study; the mechanism transfers, the agent-specific review load of untrusted input, tool access and non-determinism is outside its scope.

5. Include counsel, with a briefing artifact. Legal reviewers need architecture literacy, not only policy literacy, which means handing them a map of where decisions actually occur rather than a policy memo. That framing is a legal-trade op-ed with no adoption evidence; the underlying point holds anyway: a lawyer who cannot see where the autonomy sits cannot allocate the liability. The supervisory expectation runs the same way and is older: compliance and risk functions should be able to “understand and challenge” the algorithms, not merely receive them. What that costs, and why a CCO cannot sign without it, is compliance as an approver.

6. Cap reviews per person per week, and rotate. The cap keeps approval a decision rather than a reflex. Rotation stops one reviewer becoming the de facto owner of a build they did not write.

7. Give the pool standing to reject, and never frame the agent as an employee to them. In a randomized study of 1,261 HR and finance managers, work presented as coming from an AI employee rather than an AI tool cut reviewers’ personal accountability by nine percentage points, raised accountability attributed to the AI by eight, and increased additional review by 44%, with reviewers reporting lower confidence in their own judgment and a greater tendency to pass work onward rather than stand behind the review. Carry the scope condition: no significant effect among managers unfamiliar with AI employees, the domain was document review rather than code, and the authors have a professional stake in the anti-anthropomorphism position (see citizen developer roles).

8. Put automated checks in front of the humans. Route the mechanical checks (quality, security, policy conformance) into the pipeline so the bench spends its capacity on judgment. This is the sharpest unresolved tension on the page and it deserves the honest label: the best-known argument that human review alone gets overwhelmed by agent-generated volume is a consultancy’s, and it offers no review-throughput data, no defect-escape rate, and no volume series, so the step rests on engineering judgment, asserted, and it matches everything else here.

Where the bench comes from, in descending order of evidence: professional developers and domain experts the firm already employs; a centre of excellence staffed from whoever ran the pilot, which is Microsoft’s recommendation to its own customers with no adoption or outcome data attached; and standing review boards, which large professional-services firms demonstrably operate, though none publishes a charter, membership, cadence or decision rights, so their existence is the only transferable fact.

How you’d know it’s working

Rejection rate nonzero and stable, and no reviewer sitting at 100% approval. Per reviewer, not pool-wide: the aggregate hides the person who approves everything.

Review latency measured and bounded. Unbounded latency is how the shadow path wins; see promotion gates.

Reviewer-to-builder ratio tracked as building scales. If review load grows with citizen building and the bench does not, the constraint is tightening in real time.

Ask reviewers whether they would stand behind the review. The measured failure mode is passing work onward while calling it reviewed, and it never shows up in approval statistics.

Review is more than overhead, which is worth saying to whoever funds the bench. In a documented research deployment, an agent interprets CRISPR screens and reports confidence levels while the lab director keeps the call on which genes get an expensive follow-up experiment, and the human verification layer surfaced a result other models had dismissed as noise.

What this doesn’t solve

Human review decays under volume however well it is staffed. The pool buys a functioning gate, not immunity from alert fatigue, and volume control is an architecture problem rather than a headcount one. See oversight decay and human approval gates.

A pool cannot make a reviewer competent outside their domain, and this page’s whole argument is that domain competence is the scarce input.

One claim in circulation should be refused rather than cited: that 57% of global enterprises have introduced formal AI oversight committees or roles. The consultancy printing it attributes the figure to a 2024 report that cannot be located in any index, under a research partnership that appears not to exist. This page therefore makes no claim that standing oversight bodies are majority practice, and neither should a firm’s own policy.

Finally, the position that verification cannot be delegated to agents because accountability must remain human is a normative assertion rather than a finding. It is the right position, and it rests better on supervision obligations a regulator can enforce than on a management essay.

See also