Real citizen-developer training is measured in tens of hours (40 to 80 is the documented range), with certification depth keyed to the business criticality of what the builder will ship, and assessment graded on visible judgment, not on whether the deliverable runs.
Take the hours figure with its provenance attached: it comes from one interview program, published twice by the same two authors, in the RPA and low-code era of 2023, self-reported by the programs themselves. It is the only cross-company benchmark available, and a single source rather than two independent ones.
Start with the finding that should reshape the curriculum. In a randomized field experiment, consultants working on a task outside the AI frontier, one where the model’s confident answer is wrong, got 84.5% correct without AI. With AI they scored 70.6%. With AI plus a prompt-engineering overview they scored 60%. The trained arm did worse than the untrained arm. Those are percentage-point drops off an 84.5% base, not relative declines, and the gap between the two AI arms is significant only at the 10% level, so it needs stating carefully. The mechanism is straightforward: training that raises confidence in a tool without teaching where the tool fails is negative-value exactly where failure is expensive. See the jagged frontier.
The second finding closes the loop. In a field study of 523 early-career professionals at one firm, roughly a quarter (the authors call them AI apprentices) performed below the AI-only baseline. They would have done better by not participating. And they scored high on critical thinking, domain knowledge and AI literacy, largely indistinguishable from the group that outperformed the baseline. What separated them was application: their critiques landed on irrelevant issues, or steered the model unproductively. Possessing a skill and applying it are different capabilities. That is one unreplicated study on one firm’s task set with one agent, run by authors from that firm, so read the mechanism and not the percentages.
The checklist
1. Sequence tool literacy → failure-mode literacy → verification practice. Most curricula stop at the first. The prompt-technique layer is also the fastest-depreciating: one prominent practitioner argues those courses are already losing value, and the frontier experiment gives that argument teeth. Teach where the model breaks before teaching how to ask it nicely.
2. Teach the four foundations the research actually names. Software-development basics (lifecycle phases, version control, coding standards, the collaboration tools); relational data and data modeling, where citizen builders demonstrably struggle with how the data layer and the interface interact; troubleshooting and debugging; and basic application security, data protection and compliance. Data handling earns the opening slot: one case firm ran it as a three-to-four-day workshop before anything else.
3. Tier by the risk tier of the target build, not by job title. Certification depth keys to what the builder is allowed to ship. The widely circulated Citizen I/II/III ladder is a book’s proposal rather than an observed practice, and should not be presented as an industry standard. See risk-tier assignment.
4. Make the certificate the access gate. This is documented practice, not aspiration: 10 of 24 case companies required certified and authorized builders holding an entry ticket before platform access. In one case the training certificate is literally the credential that grants platform access. One large manufacturer required a generative-AI course, completed by more than 20,000 employees, before anyone could deploy those tools on citizen projects. That figure is company-reported, 2024 vintage, and the authors disclose a paid relationship with the firm.
5. Start from a live use case and protect the time. One energy major runs a four-month program blending instruction with hands-on coaching, entered only with a real use case not already served by the application portfolio. A professional-services firm gave its cohort time off from client work to train. Programs that ran on evenings and goodwill are the ones whose graduates regress. The often-quoted 70/20/10 split of experience to coaching to coursework is a 1980s leadership heuristic with contested evidence, repeated without citation; it works as a rule of thumb provided it is labelled as one.
6. Assess the process, not the artifact. Can the builder show why they trust this output, name what would have to be true for it to be wrong, and demonstrate the check they ran? A deliverable that runs proves nothing about the judgment behind it, and the apprentice cohort above is the evidence.
7. Pair formal training with community learning, and expect resistance. The research recommends formal foundations plus informal, experiential community learning, and names the objection any program must answer: builders routinely judge the time cost of training as not worth paying. A curriculum that cannot show the payoff stays empty; see the support model.
8. Retrain when the frontier moves. Yesterday’s inside-the-frontier task quietly becomes today’s trap when a model version changes. Also unresolved: pre-agent evidence shows citizen-developer skills are platform-bound and fail to transfer between low-code platforms. Whether prompt and agent-design skills transfer across runtimes is an open question, and transfer should not be assumed.
What actually predicts adoption
The employer, not the employee. Holding the tool constant across 66 firms, average weekly usage ranged from 6.3% at one firm to 70% at another, and firm fixed effects explained roughly twice as much variation in individual adoption as the worker’s own prior behaviour. “Firm” bundles managerial practice, licensing, job mix and training together, so this fails to isolate training, though it does say the lever sits with the organization.
Cross-country data points the same way with an important qualification. Training provision is the strongest country-level correlate of converting AI exposure into adoption (r = 0.49 across 35 countries), while individual training does not independently predict adoption once other factors are in the model. That is an ecological correlation with 35 data points, and the honest reading is the author’s own: what transfers is a standing adult-learning capability, not any particular AI course.
One encouraging result runs against the fear that rules suppress use: in a pharmaceutical field experiment, prescriptive role-specific guidelines were welcomed by the people receiving them. That reaches us secondhand through a literature review, so treat it as a lead rather than a finding.
How you’d know it’s working
Certified builders’ submissions pass promotion gates at a measurably higher rate than uncertified ones. If certification doesn’t move gate outcomes, it is theater, and it is the one measurement that would settle the question, worth instrumenting from the first cohort.
Retained capability during an outage. In the strongest longitudinal evidence available, workers kept their gains when the AI returned nothing, and the retained gain grew with months of exposure: an outage one month after adoption showed no retention, one after three months did. That held only for workers who actually followed the AI’s suggestions, making learning conditional on engagement; those who habitually deviated showed no retained gain even after prolonged access, and the authors flag that outages are rare and the estimates noisy.
What this doesn’t solve
Training doesn’t fix incentives. A builder rewarded for shipping volume will ship volume, whatever the curriculum taught.
No curriculum makes junior staff reliable verifiers of outputs beyond their own expertise; that is the reviewer pool’s problem, and the apprentice finding above is the reason it can’t be solved by teaching harder.
Two claims in this literature contradict each other and both are in the evidence: that a modest amount of training is enough for nontechnical staff to automate complex processes, and that only a small fraction of employees ever effectively master the tools. Neither this page nor the literature resolves the contradiction, so the safe course is to plan for the second and be pleased by the first.
And there is a gap under the whole field. A peer-reviewed review of citizen-developer training found firms buying programs before defining objectives, methodologies that are generic rather than tailored, and no long-term effectiveness studies at all. Everything on this page is a design argument from adjacent evidence. Nobody has measured whether any of it works over years.
See also
- The jagged frontier — why novices gain most inside the frontier and get hurt outside it.
- Citizen developer roles — who the curriculum tiers map to.
- The support model — the enablement structure the curriculum lives in.
- Risk tiers and trust zones — the criticality axis certification depth keys to.
- Promotion gates — where trained judgment gets tested.
- The reviewer pool — where the curriculum’s verification training gets used.
- Dimension 8: Training & enablement — the hub this curriculum answers to, including the survey evidence that resistance isn’t the constraint.