The problem

Staff at most firms are already building agents, the cheapest available measurement will find more of them than the inventory shows, and the citizen-development playbook that contained spreadsheets and low-code apps breaks on artifacts that act.

What changed

Non-engineers building their own software predates AI. What changed is the interface. The abstraction ladder from machine code to assembly to high-level languages added a rung: a natural-language description now compiles to a working program, so the constraint moved from who can write code to who understands the problem. In a fund, that is analysts, ops staff and PMs, none of whom report to the technology function.

The economics are measured. A field study of 5,179 customer-support workers found a 14% average productivity gain and 34% for novices, with minimal effect on the most experienced (Brynjolfsson, Li & Raymond, NBER WP 31161; GPT-3-era field data, 2020–21). A preregistered experiment on 453 professionals found tasks completed 40% faster at 18% higher quality (Noy & Zhang, Science, July 2023). Both find compression: AI helps the less expert most, which amounts to saying that the population able to build a working tool just got much larger.

Adoption followed at a rate with few precedents. By August 2024, roughly 40% of US adults aged 18–64 used generative AI and 23% of employed respondents had used it for work in the prior week. That is diffusion at least as fast as the personal computer’s, benchmarked against 1984 CPS computer-use data (Bick, Blandin & Deming, NBER WP 32966). It needs no hardware, no purchase order and no administrator, so employees bring it to work whether or not the employer participates.

Agents, new since 2025, changed what people attempt: about 23% of what users asked an agent to do were tasks they had never attempted with a conversational assistant (Perplexity researchers in HBR, July 2026; a vendor studying its own product, so read the direction, not the magnitude). OpenAI’s telemetry on its own workforce shows active Codex users up more than fivefold in the first half of 2026, growing fastest outside software engineering; trader and portfolio manager are among the job titles. It is a leading indicator of a frictionless firm rather than a description of the typical fund, as the authors say too.

The gap, and how much to trust the numbers about it

Adoption statistics here span a factor of ten, from 3% of firms in the 2019 Annual Business Survey to 88% in McKinsey’s 2025 estimate, because of sample, question wording and respondent incentive rather than disagreement about the world (NBER WP 34836 §1). Any adoption number quoted without its instrument is decoration.

The defensible version: 69% of firms across the US, UK, Germany and Australia actively use AI, from nearly 6,000 executives on central-bank-affiliated panels with verified respondent identities (NBER WP 34836, fielded Nov 2025–Jan 2026), with finance and insurance at the top of the sector distribution. One scanning vendor’s telemetry says the same about depth, with financial-services accounts averaging 119 AI components each against 79 for technology firms, so regulated industries lead on both breadth and depth.

Governance did not keep pace. Among companies planning agentic deployment, about 21% report a mature model for governing autonomous agents (Deloitte, n=3,235, fieldwork August–September 2025). About two-thirds of surveyed organizations permit employee-led AI development, 60% have a formal policy for it, and half concede low visibility into employees’ agent use (EY, October 2025); of 600 breached organizations, 63% had no AI governance policy at all (IBM/Ponemon, July 2025).

One caveat covers nearly every consultant and vendor statistic in this wiki, stated once rather than repeated: the respondents are C-suite leaders at $1B+ firms, orders of magnitude larger than a 20-to-500-person fund. Those numbers establish that the problem exists without sizing it for a firm that small.

The value paradox that makes this worse

Individual-task studies find large gains. Firm-level studies find close to nothing: more than 90% of executives report no impact on their own firm’s employment over three years and 89% none on labor productivity, while expecting +1.4% over the next three (NBER WP 34836). A randomized six-month deployment across 66 firms and 7,137 knowledge workers, and a shift-share analysis of European survey microdata, both find no change in the quantity or composition of tasks (arXiv:2504.11436; arXiv:2604.18849).

Both are sound and answer different questions, so there is no side to pick. A real individual gain that stays invisible at firm level is the ordinary shape of a general-purpose technology, and it is the condition under which ungoverned building fills the vacuum: the value is legible to the person doing the work, illegible to whoever would fund the governed version.

Why the citizen-development playbook doesn’t transfer

The pre-AI playbook of sanctioned platforms, maker certification and tiered app review assumed a reviewable artifact and a bounded blast radius. Both assumptions fail here.

There is often no artifact anybody reads. Code reaches production without any accountable human having understood it, and the builder neither can perform that review nor intends to. Integration decisions once made by engineers under review are now made by end users with none: connecting an agent to a data source or tool is a checkbox, and whoever clicks it has no reason to know they widened an attack surface.

And the artifact acts. A spreadsheet that is wrong misinforms someone; an agent that is wrong sends the email, files the ticket, or moves the position. Why that breaks design-time verification is why agents break application governance; the resulting shift from employees pasting data into chat to employees deploying applications wired into production is shadow agents.

The strongest counter-argument (100+ interviews on grassroots automation in 2024, no disaster stories in any of them) is absence of evidence from interviews, and it predates agentic deployment, when the worst case was a bad dashboard. Anyone selling agent governance should still have to explain why a decade of low-code citizen development stayed quiet.

The two axes people conflate

Three ladders get called “maturity” and mixing them produces bad decisions:

LadderRungsLives in
Implementation maturity (the builder)power user → analyst who automates → technical analyst → embedded developerhere; training curriculum
Operational criticality (the artifact)personal script → shared team tool → department workflow → business-criticalrisk tiers; tiering
Program maturity (the firm)Crawl → Walk → Run → Fly, per dimensionthe twelve dimension hubs

Plot builder discipline against artifact criticality and the failure mode names itself. Low discipline at low criticality is a sandbox: let people build, keep ceremony minimal, watch what sticks. High discipline at high criticality is governed production. The damaging quadrant is low discipline at high criticality: unreviewed code inside a critical process, no tests, no second owner, key-person risk with a cron job.

Nobody builds there deliberately; artifacts arrive by drift. A personal script gets shared, the team comes to rely on it, and criticality rises while builder discipline stays put. Governance is what lifts the artifact before the dependence grows, and that is the whole sequencing argument: watch for movement along the criticality axis, then require the controls. The two-ladder framing is this wiki’s own synthesis; the per-agent tiering it feeds is well sourced.

The rule that already applies

One correction before planning around a future AI regulation. For an SEC-registered adviser the binding rule was written in 2003, and the Commission built it to reach exactly this situation: inadequate written policies and procedures is “a violation of our rules independent of any other securities law violation,” available “before that failure has a chance to harm clients or investors.” Nothing has to go wrong. The AI-washing settlements of March 2024 charged that compliance rule alongside the marketing rule, so the theory is in active use. Meanwhile the proposals people are waiting on (adviser cybersecurity, outsourcing, predictive data analytics) were all withdrawn in June 2025, which removed specificity rather than jurisdiction; detail in regulatory exposure.

How you’d know you have this problem

Three questions to ask this week.

  1. Name every agent that can write to a production system, and its owner of record. If nobody produces that list in a day, every control the firm owns is scoped to the agents it already knows about; see agent inventory.
  2. What can an agent reach on the network, and who decided? If the answer is “whatever the builder’s laptop can reach,” there is no egress boundary and the third leg of the lethal trifecta is uncut.
  3. When someone leaves, what stops running? If that is unknown, these artifacts are already load-bearing; see offboarding.

What this page doesn’t solve

Naming the problem is not the program. Nothing here says what to build first, and the ordering matters more than the inventory of good ideas; that is sequencing, which puts architectural controls before runtime checks and says why.

Nor does this establish that agents are dangerous in any particular firm, only that adoption is real, measurement is weak, and the governance most firms have was built for a different artifact. What actually goes wrong is the concerns catalog.

See also