---
title: "Task horizon · METR's agent task-length doubling curve · Agent Fieldbook"
url: https://agentfieldbook.org/horizon/
description: "Agency measured in hours: the length of task an AI agent can complete autonomously, doubling every four to seven months per METR. Today's frontier is near a 4.8-hour module refactor. A source-graded read of the doubling curve and the capability ladder each doubling unlocks."
section: "Arena · Task Horizon"
source: Agent Fieldbook — generated from the published page
---

# Agency, measured in hours

**Agency is measured in hours.** The most useful question is not only what a model knows, but how long an agent can complete meaningful work autonomously. That duration, its task horizon, shows both the practical utility and the current limits of agent systems.

- [METR, Mar 2025 ↗](https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/)

## The capability ladder · each doubling unlocks a tier Live data

Inline completion, autocomplete-grade help

Own a single well-scoped issue end-to-end

Multi-file change across a module, CURRENT frontier

Own a ticket for a full day unsupervised

Own a whole feature across days

## Field notes

The length of task an agent can finish on its own, doubling every four to seven months, each doubling unlocking a new tier of work

**Task horizon measures how long an agent can carry a task through at a meaningful success rate.** Using a 50% success threshold on real software tasks, METR's research suggests this horizon has been doubling roughly every four to seven months. Current frontier systems can sustain work on the order of hours before a handoff becomes necessary.

If the trend holds, the next frontier is not a better answer, it is reliable execution across longer and more complex bodies of work. In that world people do not leave the loop. Their role shifts toward governance, security, approval and high-consequence judgement inside increasingly capable agent flows.

- [Brandon Chaplin](https://www.linkedin.com/in/brandon-chaplin-digital-marketing-strategist)

## Common questions

### What is the task horizon for AI agents?

The length of task an agent can finish on its own, at a fifty per cent success rate. It measures stamina rather than knowledge: how far an agent gets before it needs a human. The research group METR introduced the measure using real software tasks, which is why it is read in minutes and hours rather than in benchmark scores.

### How fast is the task horizon growing?

It has been doubling roughly every four to seven months, on METR's measurements. Plotted on a log scale that is close to a straight line, from about a minute in 2019 to several hours today. Anything past the measured points is an extrapolation of that doubling rate, not an observation.

### How long can an AI agent work without human help?

The frontier is currently around five hours on a well-scoped engineering task, such as refactoring a module. Shorter work, like a code snippet or a single bug fix, is comfortably inside that. A full working day or a multi-day project is not there yet.

### Why does each doubling matter?

Because work comes in units of time. Once an agent can run for longer than a task takes, it can own the whole task instead of a person stitching the steps together. Each doubling moves that line up a rung: snippet, bug fix, working day, project.

### Will AI agents replace human workers?

Not on the current evidence. A longer horizon moves more of the doing to agents, but someone still has to set the goal, approve the risky steps and answer for the result. The realistic read is that roles shift towards direction, review and accountability.

## Next in the learning path

- [Agent mobility The web tipping from humans to agents](https://agentfieldbook.org/mobility/)

- [Org models How agentic teams are structured](https://agentfieldbook.org/workforce/org_models/)

- [Productivity Measured agent productivity gains](https://agentfieldbook.org/workforce/productivity/)

- [Benchmarks The evaluations behind the numbers](https://agentfieldbook.org/software/benchmarks/)

- [All agents The systems pushing the frontier](https://agentfieldbook.org/all_agents/)

## Entries

_5 entries listed on this page._

| name | description | subjectOf |
| --- | --- | --- |
| Code snippet | Inline completion, autocomplete-grade help | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ |
| Bug fix | Own a single well-scoped issue end-to-end | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ |
| Module refactor | Multi-file change across a module, CURRENT frontier | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ |
| Full workday | Own a ticket for a full day unsupervised | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ |
| Multi-day project | Own a whole feature across days | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ |
