Corpus Agentis
The field book to agent ecosystems
The field book to agent ecosystems
Arena · Coming soon

Head-to-head, task-by-task

Two- and three-agent side-by-side comparisons, ELO-ranked head-to-heads, and capability-by-task matrices: the buyer's comparison view. In build; the measured views it draws on are already live.

In buildUntil Arena ships, the measured views are already live: Stack → Benchmarks and Leaderboard for capability scores, Overview → Task Horizon for the agency-in-hours doubling curve, and any agent card for its full spec sheet.
Common questions
What will the Arena do?

Put agents head to head, task by task. Side-by-side comparisons, ELO-ranked matchups and a capability-by-task matrix. The question it answers is which agent to use for a specific job, not which one is best overall.

Is the Arena live yet?

No, it is still in build. The evidence behind it is already live: Benchmarks and Leaderboard for capability scores, Task Horizon for how long agents can run unaided, and every agent card for its full sourced spec.

How do I compare AI agents?

Start with the task, not the brand. Check benchmark scores for the kind of work you actually need, then how long the agent can run before it needs help, then the practical things: which tools it connects to, where it runs and what it costs at your volume. Two agents rarely win on the same job.

Next in the learning path