Corpus Agentis
The field book to agent ecosystems
The field book to agent ecosystems
Digital Workforce · Productivity

Productivity, measured

How much an agent fleet or copilot actually changes human output, taken from measured studies rather than vendor promises. The benchmarks below are randomised trials, field experiments and production deployments, each graded by cohort and confidence, and they keep the counterpoints in view where experienced workers got slower.

Illustrative model, not a forecast. It multiplies out a fleet at the throughput and success rate you set, against a measured human baseline. The point is the shape of the trade, not the exact number.

Model: augmented output = human baseline + (agents per human × agent throughput × success rate). FTE-equivalent = augmented ÷ human baseline. The unit view multiplies by team size. Tied to the task-horizon trend: as agents handle longer tasks, throughput climbs.

The evidence at a glance
-20%0%+20%+40%+60%Baseline · human aloneGitHub Copilot RCT95 devs, controlled+55%GitHub Copilot pooled3 RCTs, n≈4,867+26%BCG × Harvard758 consultants, GPT-4+25%NBER Generative AI at Work5,000+ CS agents, RCT+14%METR 2025 (counterpoint)16 experienced OSS devs-19%Productivity delta vs baseline (percentage points)STUDY · COHORT

Klarna's ≈700-FTE deployment and the McKinsey adoption survey use different axes and stay in the cards below. The METR counterpoint is plotted on the negative side to keep the "not universal" reading in view.

Real-world productivity benchmarks · measured, sourced Live data
Customer Support
≈700 FTE / 1 AI system

Klarna's OpenAI-powered AI assistant did the work of ~700 full-time agents: 2.3M conversations (⅔ of CS chats) in month one; resolution 11min → <2min; 25% fewer repeat inquiries; $40M profit lift (2024).

Customer Support
+14% avg · +34% novice

NBER "Generative AI at Work" RCT: AI assistance raised issues resolved per hour by 14% on average: and 34% for the least-experienced workers.

Software Engineering
+55% task speed

GitHub Copilot controlled study (95 devs): the AI group finished a task 55% faster (95% CI 21–89%). Largest gains for less-experienced developers.

Software Engineering
−19% (experienced)

COUNTERPOINT, METR 2025 RCT (16 experienced OSS devs): with AI they were 19% SLOWER on complex familiar tasks, while believing they were 20% faster. Gains are not universal.

Knowledge Work / Consulting
+12.2% tasks / +25.1% speed / +40% quality

Field experiment with 758 BCG consultants: GPT-4 users completed 12.2% more tasks, worked 25.1% faster, +40% quality on tasks inside the AI frontier; but 19pp less likely to be correct on tasks outside it (the jagged frontier).

Software Engineering
+26.08% completed tasks (n≈4,867, 3 RCTs)

Three RCTs (Microsoft, Accenture, an anonymous Fortune 100 company), pooled n≈4,867: GitHub Copilot raised completed tasks by ~26.08%; less-experienced developers gained most.

Enterprise AI Adoption
23% scaling agents / 39% piloting

McKinsey State of AI (Nov 2025, 1,993 respondents): ~80% use genAI in ≥1 function, 23% scaling an agentic system somewhere, 39% experimenting; in any single function ≤10% say scaling agents; ~5.5% report significant EBIT impact.

Field notes

What the RCTs, field studies and real deployments actually found

The studies do not agree. A support trial found 14% more resolutions per hour, rising to 34% for the least experienced staff. A coding study cut task time by 55%. A 2025 trial of experienced developers found the opposite: 19% slower, while the developers believed they were faster.

Show more

The split is consistent. AI helps most where the worker is new and the task is routine. It helps least, and can hurt, where the person already knows the work. A headline number only describes the group and the task it was measured on.

From the corpus, curated by Brandon Chaplin
Common questions
How is the fleet multiple calculated?

Augmented output = the human baseline + (agents per human x tasks per agent per day x success rate). The multiple is that total divided by the baseline. Business-unit figures are the same maths scaled by team size. It is deliberately simple, so you can see which input is driving the answer.

Where does the human baseline come from?

Published salary and task-volume benchmarks per role, carried in the corpus with a source on each row. The baseline is tasks per day for that role, not an estimate of effort, so the comparison stays like for like.

What is the success rate slider doing?

Discounting agent output. At 80% success, only 80% of attempted tasks count toward the total, which is why the multiple falls sharply as reliability drops. It is the single input the result is most sensitive to, alongside fleet size.

Are the augmentation percentages measured or modelled?

Measured. The uplift figure shown for each role comes from a published study with its cohort and source linked. Where no controlled trial exists for a role, the tool says so rather than filling the gap with an estimate.

Why does this disagree with the benchmarks below it?

Because they answer different questions. The calculator multiplies out a fleet doing task work. The benchmarks measure humans working alongside AI, which includes cases where experienced staff got slower. Read the benchmarks as the evidence and the calculator as the arithmetic.

Next in the learning path