Agency, measured in hours
Agency is measured in hours. The most useful question is not only what a model knows, but how long an agent can complete meaningful work autonomously. That duration, its task horizon, shows both the practical utility and the current limits of agent systems.
Code snippet
Inline completion, autocomplete-grade help
Bug fix
Own a single well-scoped issue end-to-end
Module refactor
Multi-file change across a module, CURRENT frontier
Full workday
Own a ticket for a full day unsupervised
Multi-day project
Own a whole feature across days
The length of task an agent can finish on its own, doubling every four to seven months, each doubling unlocking a new tier of work
Task horizon measures how long an agent can carry a task through at a meaningful success rate. Using a 50% success threshold on real software tasks, METR's research suggests this horizon has been doubling roughly every four to seven months. Current frontier systems can sustain work on the order of hours before a handoff becomes necessary.
Show more
If the trend holds, the next frontier is not a better answer, it is reliable execution across longer and more complex bodies of work. In that world people do not leave the loop. Their role shifts toward governance, security, approval and high-consequence judgement inside increasingly capable agent flows.
What is the task horizon for AI agents?
The length of task an agent can finish on its own, at a fifty per cent success rate. It measures stamina rather than knowledge: how far an agent gets before it needs a human. The research group METR introduced the measure using real software tasks, which is why it is read in minutes and hours rather than in benchmark scores.
How fast is the task horizon growing?
It has been doubling roughly every four to seven months, on METR's measurements. Plotted on a log scale that is close to a straight line, from about a minute in 2019 to several hours today. Anything past the measured points is an extrapolation of that doubling rate, not an observation.
How long can an AI agent work without human help?
The frontier is currently around five hours on a well-scoped engineering task, such as refactoring a module. Shorter work, like a code snippet or a single bug fix, is comfortably inside that. A full working day or a multi-day project is not there yet.
Why does each doubling matter?
Because work comes in units of time. Once an agent can run for longer than a task takes, it can own the whole task instead of a person stitching the steps together. Each doubling moves that line up a rung: snippet, bug fix, working day, project.
Will AI agents replace human workers?
Not on the current evidence. A longer horizon moves more of the doing to agents, but someone still has to set the goal, approve the risky steps and answer for the result. The realistic read is that roles shift towards direction, review and accountability.