Corpus Agentis
The field book to agent ecosystems
The field book to agent ecosystems
Arena · Task Horizon

Agency, measured in hours

Agency is measured in hours. The most useful question is not only what a model knows, but how long an agent can complete meaningful work autonomously. That duration, its task horizon, shows both the practical utility and the current limits of agent systems.

1min10min1hr8hr17hr40hr20192020202120222023202420252026202720284.8 hours: Module refactor (current)2026-01 · 4.8hr2026-08 · 7.666666666666667hr2027-01 · 10.666666666666666hr2027-07 · 17hr2027-10 · 25hr2028-01 · 38.333333333333336hr2019-05 · 1min2020-05 · 1.8min2021-05 · 2.8min2022-07 · 5min2023-05 · 8.5min2023-11 · 11min2024-05 · 15min2024-11 · 26min2025-03 · 38min2025-08 · 1.2hr2025-11 · 2.8hr2026-01 · 4.8hr Task horizon (50% success, log)
Measured (METR)   Projected, 4–7mo doubling (estimate)  ·  METR, Mar 2025 ↗
The capability ladder · each doubling unlocks a tier Live data
seconds–minutes

Code snippet

Inline completion, autocomplete-grade help

passedconfirmed
1 hour

Bug fix

Own a single well-scoped issue end-to-end

passedconfirmed
4.8 hours · current frontier

Module refactor

Multi-file change across a module, CURRENT frontier

currentconfirmed
8 hours · projected

Full workday

Own a ticket for a full day unsupervised

projectedprojected
32 hours · projected

Multi-day project

Own a whole feature across days

projectedprojected
Field notes

The length of task an agent can finish on its own, doubling every four to seven months, each doubling unlocking a new tier of work

Task horizon measures how long an agent can carry a task through at a meaningful success rate. Using a 50% success threshold on real software tasks, METR's research suggests this horizon has been doubling roughly every four to seven months. Current frontier systems can sustain work on the order of hours before a handoff becomes necessary.

Show more

If the trend holds, the next frontier is not a better answer, it is reliable execution across longer and more complex bodies of work. In that world people do not leave the loop. Their role shifts toward governance, security, approval and high-consequence judgement inside increasingly capable agent flows.

From the corpus, curated by Brandon Chaplin
Common questions
What is the task horizon for AI agents?

The length of task an agent can finish on its own, at a fifty per cent success rate. It measures stamina rather than knowledge: how far an agent gets before it needs a human. The research group METR introduced the measure using real software tasks, which is why it is read in minutes and hours rather than in benchmark scores.

How fast is the task horizon growing?

It has been doubling roughly every four to seven months, on METR's measurements. Plotted on a log scale that is close to a straight line, from about a minute in 2019 to several hours today. Anything past the measured points is an extrapolation of that doubling rate, not an observation.

How long can an AI agent work without human help?

The frontier is currently around five hours on a well-scoped engineering task, such as refactoring a module. Shorter work, like a code snippet or a single bug fix, is comfortably inside that. A full working day or a multi-day project is not there yet.

Why does each doubling matter?

Because work comes in units of time. Once an agent can run for longer than a task takes, it can own the whole task instead of a person stitching the steps together. Each doubling moves that line up a rung: snippet, bug fix, working day, project.

Will AI agents replace human workers?

Not on the current evidence. A longer horizon moves more of the doing to agents, but someone still has to set the goal, approve the risky steps and answer for the result. The realistic read is that roles shift towards direction, review and accountability.

Next in the learning path