Corpus Agentis
The field book to agent ecosystems
The field book to agent ecosystems
Arena · Capability leaderboards

Top agents by capability

Real benchmark results across the capabilities agents are hired for: coding, science/reasoning, math, factuality. Epoch AI (CC-BY) + Aider Polyglot (Apache-2.0), live. Pick a capability.

Aider Polyglot CodingGPQA Diamond Science / ReasoningAIME MathFrontierMath Frontier MathSimpleQA FactualitySWE-bench Verified CodingMMMU MultimodalOSWorld Computer UseArtificial Analysis Image Arena ImageArtificial Analysis Video Arena VideoArtificial Analysis Speech Arena AudioArtificial Analysis Output Speed Speed

Aider Polyglot, Multi-language code editing across a polyglot repository.

Coding · Aider Polyglot Live data
1.gpt-5 (high)88%
2.gpt-5 (medium)86.7%
3.o3-pro (high)84.9%
4.gemini-2.5-pro-preview-06-05 (32k think)83.1%
5.o3 (high)81.3%
6.gpt-5 (low)81.3%
7.grok-4 (high)79.6%
8.gemini-2.5-pro-preview-06-05 (default think)79.1%
9.o3 (high) + gpt-4.178.2%
10.Gemini 2.5 Pro Preview 05-0676.9%
11.o376.9%
12.DeepSeek-V3.2-Exp (Reasoner)74.2%
13.Gemini 2.5 Pro Preview 03-2572.9%
14.o4-mini (high)72%
15.claude-opus-4-20250514 (32k thinking)72%
Bar baseline zoomed to 65.9% so ranking reads · numbers are the true values
Models behind these scores · in the supply-chain network
OpenAI GPT · ~25% enterprise prod share (Menlo, mid-2025)Anthropic Claude · ~32% enterprise prod share, ahead of OpenAI (Menlo, mid-2025)Google GeminiMeta Llama · OSS leaderMistral · EU altDeepSeek · cheapest frontierxAI Grok · Grok 4.x frontier modelsAlibaba Qwen · Open-weight Qwen3 seriesAmazon Nova · AWS Nova FM family (Bedrock)Cohere · Command enterprise LLMs (RAG / agents)
Field notes

Real benchmark results across the capabilities agents are hired for

Leaderboards rank models on the capabilities agents draw from, and the spread at the top has narrowed sharply. On most tasks the leading models are separated by a point or two, well inside the noise of any single evaluation.

From the corpus, curated by Brandon Chaplin
Common questions
What is an AI model leaderboard?

A table ranking models by their score on one benchmark. The same fixed task set is run against each model and the results are sorted. Because each leaderboard covers a single capability, a model that tops coding can sit mid-table on maths.

Which AI model is best for coding?

It depends on the coding work. For fixing real issues in an existing codebase, look at SWE-bench Verified; for multi-language editing, the Aider polyglot leaderboard. The leading models usually sit within a couple of points of each other, so price and latency often decide it. (Aider)

What is GPQA?

A set of multiple-choice science questions written by PhD holders in biology, physics and chemistry. They are hard enough that skilled non-experts with unrestricted web access score around 34%. It is used to test graduate-level reasoning rather than recall. (GPQA paper)

Why do the top AI models score so close together?

Because the benchmarks are close to saturation. Once several models sit within a point or two, the gap is inside the noise of a single evaluation run and the ranking stops picking a winner. Selection then moves to price, latency, context length and tool-calling reliability.

Who publishes AI benchmark scores?

Some come from the labs themselves, which is why independent runs matter. Epoch AI publishes its own evaluations under CC-BY, and the Aider polyglot leaderboard is published under Apache 2.0. New models are scored as they ship, so any ranking is a snapshot. (Epoch AI)

Can AI labs game benchmark results?

Mostly by accident. Benchmark questions leak into training data over time, which inflates scores without improving the model. Labs also choose which benchmarks to report and which prompting setup to use. Independent runs and held-out test sets are the usual defence.

Next in the learning path