---
title: "AI agent productivity, measured · RCTs, field studies, real deployments · Agent Fieldbook"
url: https://agentfieldbook.org/workforce/productivity/
description: "What agents and copilots actually do to human output, taken from measured studies rather than vendor promises: Klarna, the NBER support RCT, GitHub Copilot trials, the METR counterpoint and the BCG jagged-frontier experiment. Each benchmark is graded by cohort and confidence, with the primary source linked, including the cases where experienced workers got slower."
section: "Digital Workforce · Productivity"
source: Agent Fieldbook — generated from the published page
---

# Productivity, measured

**≈700 FTE / 1 AI system · +14% avg · +34% novice · +55% task speed · −19% (experienced) · +12.2% tasks / +25.1% speed / +40% quality · +26.08% completed tasks (n≈4,867, 3 RCTs) · 23% scaling agents / 39% piloting**

How much an agent fleet or copilot actually changes human output, taken from measured studies rather than vendor promises. The benchmarks below are randomised trials, field experiments and production deployments, each graded by cohort and confidence, and they keep the counterpoints in view where experienced workers got slower.

Illustrative model, not a forecast. It multiplies out a fleet at the throughput and success rate you set, against a measured human baseline. The point is the shape of the trade, not the exact number.

**Model:** augmented output = human baseline + (agents per human × agent throughput × success rate). FTE-equivalent = augmented &divide; human baseline. The unit view multiplies by team size. Tied to the [task-horizon](https://agentfieldbook.org/horizon/) trend: as agents handle longer tasks, throughput climbs.

## The evidence at a glance

Klarna's ≈700-FTE deployment and the McKinsey adoption survey use different axes and stay in the cards below. The METR counterpoint is plotted on the negative side to keep the "not universal" reading in view.

## Real-world productivity benchmarks · measured, sourced Live data

Klarna's OpenAI-powered AI assistant did the work of ~700 full-time agents: 2.3M conversations (⅔ of CS chats) in month one; resolution 11min → <2min; 25% fewer repeat inquiries; $40M profit lift (2024).

- [↗ production · confirmed](https://www.klarna.com/international/press/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month/)

NBER "Generative AI at Work" RCT: AI assistance raised issues resolved per hour by 14% on average: and 34% for the least-experienced workers.

- [↗ 5,000+ agents, RCT · confirmed](https://www.nber.org/papers/w31161)

GitHub Copilot controlled study (95 devs): the AI group finished a task 55% faster (95% CI 21–89%). Largest gains for less-experienced developers.

- [↗ 95 devs, controlled · confirmed](https://github.blog/news-insights/research/research-quantifying-github-copilots-impact-on-developer-productivity-and-happiness/)

COUNTERPOINT, METR 2025 RCT (16 experienced OSS devs): with AI they were 19% SLOWER on complex familiar tasks, while believing they were 20% faster. Gains are not universal.

- [↗ 16 expert devs, RCT · confirmed](https://arxiv.org/abs/2507.09089)

Field experiment with 758 BCG consultants: GPT-4 users completed 12.2% more tasks, worked 25.1% faster, +40% quality on tasks inside the AI frontier; but 19pp less likely to be correct on tasks outside it (the jagged frontier).

- [↗ field-experiment · confirmed](https://www.hbs.edu/faculty/Pages/item.aspx?num=64700)

Three RCTs (Microsoft, Accenture, an anonymous Fortune 100 company), pooled n≈4,867: GitHub Copilot raised completed tasks by ~26.08%; less-experienced developers gained most.

- [↗ RCT · confirmed](https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf)

McKinsey State of AI (Nov 2025, 1,993 respondents): ~80% use genAI in ≥1 function, 23% scaling an agentic system somewhere, 39% experimenting; in any single function ≤10% say scaling agents; ~5.5% report significant EBIT impact.

- [↗ survey · confirmed](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai)

- [Next · Cost calculator Turn these deltas into a human-vs-agentic team cost Pick a role, set team size and automatable share, then price the agent side against live per-event vendor rates. →](https://agentfieldbook.org/workforce/calculator/)

## Field notes

What the RCTs, field studies and real deployments actually found

The studies do not agree. A support trial found 14% more resolutions per hour, rising to 34% for the least experienced staff. A coding study cut task time by 55%. A 2025 trial of experienced developers found the opposite: 19% slower, while the developers believed they were faster.

The split is consistent. AI helps most where the worker is new and the task is routine. It helps least, and can hurt, where the person already knows the work. A headline number only describes the group and the task it was measured on.

- [Brandon Chaplin](https://www.linkedin.com/in/brandon-chaplin-digital-marketing-strategist)

## Common questions

### How is the fleet multiple calculated?

Augmented output = the human baseline + (agents per human x tasks per agent per day x success rate). The multiple is that total divided by the baseline. Business-unit figures are the same maths scaled by team size. It is deliberately simple, so you can see which input is driving the answer.

### Where does the human baseline come from?

Published salary and task-volume benchmarks per role, carried in the corpus with a source on each row. The baseline is tasks per day for that role, not an estimate of effort, so the comparison stays like for like.

### What is the success rate slider doing?

Discounting agent output. At 80% success, only 80% of attempted tasks count toward the total, which is why the multiple falls sharply as reliability drops. It is the single input the result is most sensitive to, alongside fleet size.

### Are the augmentation percentages measured or modelled?

Measured. The uplift figure shown for each role comes from a published study with its cohort and source linked. Where no controlled trial exists for a role, the tool says so rather than filling the gap with an estimate.

### Why does this disagree with the benchmarks below it?

Because they answer different questions. The calculator multiplies out a fleet doing task work. The benchmarks measure humans working alongside AI, which includes cases where experienced staff got slower. Read the benchmarks as the evidence and the calculator as the arithmetic.

## Next in the learning path

- [Org models The org shapes these gains reshape](https://agentfieldbook.org/workforce/org_models/)

- [Architectures Real systems that produce these numbers](https://agentfieldbook.org/design/architectures/)

- [Human oversight The review that keeps the gains safe](https://agentfieldbook.org/security/oversight/)

- [Agent memory What lets a fleet sustain output](https://agentfieldbook.org/memory/types/)

## Entries

_7 entries listed on this page._

| name | description | url |
| --- | --- | --- |
| Customer Support · ≈700 FTE / 1 AI system | Klarna's OpenAI-powered AI assistant did the work of ~700 full-time agents: 2.3M conversations (⅔ of CS chats) in month one; resolution 11min → <2min; 25% fewer repeat inquiries; $40M profit lift (2024). | https://www.klarna.com/international/press/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month/ |
| Customer Support · +14% avg · +34% novice | NBER "Generative AI at Work" RCT: AI assistance raised issues resolved per hour by 14% on average: and 34% for the least-experienced workers. | https://www.nber.org/papers/w31161 |
| Software Engineering · +55% task speed | GitHub Copilot controlled study (95 devs): the AI group finished a task 55% faster (95% CI 21–89%). Largest gains for less-experienced developers. | https://github.blog/news-insights/research/research-quantifying-github-copilots-impact-on-developer-productivity-and-happiness/ |
| Software Engineering · −19% (experienced) | COUNTERPOINT, METR 2025 RCT (16 experienced OSS devs): with AI they were 19% SLOWER on complex familiar tasks, while believing they were 20% faster. Gains are not universal. | https://arxiv.org/abs/2507.09089 |
| Knowledge Work / Consulting · +12.2% tasks / +25.1% speed / +40% quality | Field experiment with 758 BCG consultants: GPT-4 users completed 12.2% more tasks, worked 25.1% faster, +40% quality on tasks inside the AI frontier; but 19pp less likely to be correct on tasks outside it (the jagged frontier). | https://www.hbs.edu/faculty/Pages/item.aspx?num=64700 |
| Software Engineering · +26.08% completed tasks (n≈4,867, 3 RCTs) | Three RCTs (Microsoft, Accenture, an anonymous Fortune 100 company), pooled n≈4,867: GitHub Copilot raised completed tasks by ~26.08%; less-experienced developers gained most. | https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf |
| Enterprise AI Adoption · 23% scaling agents / 39% piloting | McKinsey State of AI (Nov 2025, 1,993 respondents): ~80% use genAI in ≥1 function, 23% scaling an agentic system somewhere, 39% experimenting; in any single function ≤10% say scaling agents; ~5.5% report significant EBIT impact. | https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai |
