Capability is jagged, not level. Two tasks that look equally hard sit on opposite sides of a boundary the model never announces, and just outside it the tool makes competent people measurably worse. The governing constraint is whether a firm’s people can tell which side of the line they are on, not how accurate the model is on average.
The mechanism
The canonical experiment is a preregistered field study with 758 BCG consultants, roughly 7% of that firm’s individual-contributor consulting workforce, all using GPT-4 as it existed at the end of April 2023 with default system prompts. Date that; it matters, and the paper’s 2026 publication date misleads.
Inside the frontier the results are what the marketing promises: consultants with AI completed 12.2% more tasks, 25.1% faster, at significantly higher quality. The quality gain ran about 30–34% over a control mean, depending on whether participants also received an overview of the tool.
Outside it, the same access made people worse. On a task designed to fall just beyond the model’s competence, AI-enabled participants were 19 percentage points less likely to produce a correct answer than a control group that was right 84.5% of the time. Nothing signalled the change. The tasks looked comparable, the output remained fluent, and the authors note that AI-assisted answers were more coherent and persuasive whether or not they were right.
The shape has a name because the shape is the point. Similar-looking tasks fall on either side of the boundary, and it is hard for a knowledge worker to grasp ex ante where the boundary sits at any given moment. Karpathy’s version is blunter: the same system behaves like a genius polymath on one request and a confused schoolchild on the next.
Two structural consequences follow. First, one out-of-frontier task inside an otherwise inside-frontier workflow disproportionately drags overall performance. A five-step agent workflow is only as good as its worst-matched step, and averaging across the workflow hides it. Second, the frontier is jagged but not fixed. A task can move inside with a new model release, which means any map of what a firm’s agents are good at is dated the moment it is written.
The distributional pattern is consistent across independent studies and is the most useful thing here for staffing. Gains concentrate among the less experienced: 15% average productivity improvement across 5,172 customer-support agents but around 30% for the least experienced; 26% average developer output gain with 27–39% for short-tenure engineers against 8–13% for long-tenure ones. Noy and Zhang found the correlation between baseline ability and output falling from 0.49 in control to 0.25 in treatment. The tool compresses the distribution: it raises the floor more than the ceiling. And in the support-agent study the most skilled tercile showed a small but statistically significant quality decline.
The Cybernetic Teammate experiment (preregistered, 2×2, 776 P&G professionals on real innovation challenges) adds the finding that reframes team design. Individuals with AI matched the performance of teams without AI. Teams without AI beat individuals without AI by 0.24 SD; individuals with AI gained 0.37 SD and teams with AI 0.39 SD. The difference between team-with-AI and team-without-AI fell short of statistical significance. Time fell 12–16%. Teams with AI were 9.2 percentage points more likely to land in the top decile against a 5.8% control base, roughly triple.
The unsettling half of that study is the calibration result: AI-enabled participants were 9.2 points less likely to expect a top-10% ranking, significantly so. They performed better and believed they had done worse. Whatever internal signal tells a professional “that went well” stops tracking actual quality once a model is involved. Seen from the inside, that is the same problem the frontier poses.
Kolt supplies the framing that makes this a governance concept rather than a productivity one. Human-agent harm is a motivation problem: an employee’s interests diverge from the firm’s, so the firm aligns incentives and monitors. AI-agent harm is a competence problem: the agent fails out-of-distribution, with no divergent interest anywhere. Incentive design does nothing about it. That is why frontier position belongs in the risk model and not just the training deck.
What closes the gap is context, and it closes it only partly. Retrieval over a firm’s own corpus improves performance on proprietary codebases and internal languages without eliminating the gap, and 11 of 12 companies in one practitioner study named context management as their primary technical bottleneck. Users say the same from the other side: the most common reason people don’t use a model at work is that it lacks context-specific knowledge. Meanwhile users self-sort by task type: roughly 70% prefer AI for drafting email and 65% for basic analysis, while humans are preferred by around nine to one for complex or long-term work.
Finally, a caution about vendor numbers. A vendor reporting 98% “resolution” on KYC alerts against 55% on sanctions screening is describing alerts closed without human escalation, not accuracy against ground truth, with no sample size or independent evaluation disclosed. That gap is a decent illustration of jaggedness and nothing like a benchmark. Macro forecasts disagree by roughly twentyfold on economy-wide productivity effects, which shows how little anyone knows at the aggregate level even where the task-level results are solid.
What to do
Have each builder find the edge on their own task before trusting the output. Run the same agent against ten cases where the right answer is already known, including two the builder expects it to get wrong, and look at where it fails rather than how often. That exercise costs an afternoon and produces something no vendor benchmark can: a frontier map for the firm’s own work.
Decompose workflows and locate each step relative to the frontier separately. Because one bad step drags the whole chain, the useful question is never “can the agent do this job” but “which step in this job is nearest the edge.”
Sort builders by behaviour rather than credentials. Foundational skills (critical thinking, AI literacy, domain knowledge) failed to predict who worked well with AI; working style did, and the tool magnifies differences in how people apply skills rather than levelling them; see citizen developer roles.
Staff for the compression finding. If gains concentrate among the least experienced people and the most experienced may lose a little quality, the deployment decision is per-role, not per-firm, and senior staff are best used checking the frontier rather than working inside it.
Re-run the map on model changes. The frontier moves; a pinned model snapshot makes that a scheduled event rather than a surprise. See evals.
How you’d know it’s working
Someone can name a task the firm’s agents are bad at, specifically, and explain how they found out. A team that cannot name one has not looked.
Error-catch rate on seeded near-frontier cases rises with training. That is the direct measure of whether frontier awareness is real or rhetorical, and it belongs in training.
Gains show up in a firm metric someone else owns, not only in self-reports. Self-assessment is demonstrably miscalibrated in exactly this setting.
What this doesn’t solve
Every number here is dated, and most of it is 2023-vintage GPT-4. The shape (jaggedness, compression, miscalibration) replicates across independent studies and different tools. The magnitudes should not be quoted as current. Note also that the jagged-frontier framing and the specific task descriptions fell outside the original preregistration, which makes them the authors’ interpretation of a preregistered result rather than a preregistered hypothesis.
Frontier awareness does nothing against adversaries. A perfectly calibrated user who knows exactly where the model is weak is still injectable, and the agent is still over-privileged. See prompt injection.
Homogenization is a real cost, and this page leaves it unsolved. AI-aided solutions in the Cybernetic Teammate study were semantically more similar to each other than human-only ones. Where a whole research team works inside the same frontier with the same model, its outputs should be expected to converge, and that is a portfolio problem before it is a governance one.
And the scope condition on the headline experiment matters: its results apply where quality within an allocated time frame is prioritized over raw speed. A workflow optimizing for throughput at tolerable quality is a different regime, and the trade there has not been measured.
See also
- Unreliable output acting on your systems — what outside-the-frontier failure does once an agent acts on it.
- Oversight decay — why trust earned inside the frontier carries over to ground where it was never warranted.
- Training curriculum — where frontier-probing gets taught.
- Evals — the standing test for where the frontier actually is this month.
- Citizen developer roles — why working style, not title, sorts builders.
- Risk tiers — frontier position as an input to how much governance a use case earns.