---
title: "Agent capability leaderboards · models by benchmark · Agent Fieldbook"
url: https://agentfieldbook.org/software/leaderboard/
description: "Real benchmark results ranking AI models by capability: coding, science and reasoning, mathematics and factuality. Provider-coloured bars from Epoch AI (CC-BY) and Aider Polyglot (Apache 2.0), each score graded and linked to its source."
section: "Arena · Capability leaderboards"
source: Agent Fieldbook — generated from the published page
---

# Top agents by capability

**OpenAI GPT · ~25% enterprise prod share (Menlo, mid-2025) · Anthropic Claude · ~32% enterprise prod share, ahead of OpenAI (Menlo, mid-2025) · Google Gemini · Meta Llama · OSS leader · Mistral · EU alt · DeepSeek · cheapest frontier · xAI Grok · Grok 4.x frontier models · Alibaba Qwen · Open-weight Qwen3 series · Amazon Nova · AWS Nova FM family (Bedrock) · Cohere · Command enterprise LLMs (RAG / agents)**

Real benchmark results across the capabilities agents are hired for: coding, science/reasoning, math, factuality. Epoch AI (CC-BY) + Aider Polyglot (Apache-2.0), live. Pick a capability.

**Aider Polyglot**, Multi-language code editing across a polyglot repository.

## Coding · Aider Polyglot Live data

## Bar baseline zoomed to 65.9% so ranking reads · numbers are the true values

**GPQA Diamond**, Graduate-level science Q&A (biology, physics, chemistry), "Google-proof" questions experts get right and laypeople can't.

## Science / Reasoning · GPQA Diamond Live data

## Bar baseline zoomed to 87.6% so ranking reads · numbers are the true values

**AIME**, Competition mathematics, the American Invitational Mathematics Examination.

## Math · AIME Live data

## Bar baseline zoomed to 91.3% so ranking reads · numbers are the true values

**FrontierMath**, Research-level mathematics, the hardest math benchmark, problems that take specialists hours.

## Frontier Math · FrontierMath Live data

## Bar baseline zoomed to 45.5% so ranking reads · numbers are the true values

**SimpleQA**, Short-fact accuracy, does the model hallucinate on simple factual questions?

## Factuality · SimpleQA Live data

## Bar baseline zoomed to 40.8% so ranking reads · numbers are the true values

**SWE-bench Verified**, Can it fix real GitHub issues? A human-verified subset of resolved pull requests, the standard for agentic coding.

## Coding · SWE-bench Verified Live data

## Bar baseline zoomed to 69.9% so ranking reads · numbers are the true values

**MMMU**, Multimodal college-level reasoning across images, diagrams and text.

## Multimodal · MMMU Live data

## Bar baseline zoomed to 78.4% so ranking reads · numbers are the true values

**OSWorld**, Real computer-use tasks in a live operating system, files, browser, apps.

## Computer Use · OSWorld Live data

## Bar baseline zoomed to 60.9% so ranking reads · numbers are the true values

## Image · Artificial Analysis Image Arena Live data

## Bar baseline zoomed to 1107.7% so ranking reads · numbers are the true values

## Video · Artificial Analysis Video Arena Live data

## Bar baseline zoomed to 909.0% so ranking reads · numbers are the true values

## Audio · Artificial Analysis Speech Arena Live data

## Bar baseline zoomed to 989.1% so ranking reads · numbers are the true values

## Speed · Artificial Analysis Output Speed Live data

## Models behind these scores · in the supply-chain network

## Field notes

Real benchmark results across the capabilities agents are hired for

Leaderboards rank models on the capabilities agents draw from, and the spread at the top has narrowed sharply. On most tasks the leading models are separated by a point or two, well inside the noise of any single evaluation.

- [Brandon Chaplin](https://www.linkedin.com/in/brandon-chaplin-digital-marketing-strategist)

## Common questions

### What is an AI model leaderboard?

A table ranking models by their score on one benchmark. The same fixed task set is run against each model and the results are sorted. Because each leaderboard covers a single capability, a model that tops coding can sit mid-table on maths.

### Which AI model is best for coding?

It depends on the coding work. For fixing real issues in an existing codebase, look at SWE-bench Verified; for multi-language editing, the Aider polyglot leaderboard. The leading models usually sit within a couple of points of each other, so price and latency often decide it. [(Aider)](https://aider.chat/docs/leaderboards/)

### What is GPQA?

A set of multiple-choice science questions written by PhD holders in biology, physics and chemistry. They are hard enough that skilled non-experts with unrestricted web access score around 34%. It is used to test graduate-level reasoning rather than recall. [(GPQA paper)](https://arxiv.org/abs/2311.12022)

### Why do the top AI models score so close together?

Because the benchmarks are close to saturation. Once several models sit within a point or two, the gap is inside the noise of a single evaluation run and the ranking stops picking a winner. Selection then moves to price, latency, context length and tool-calling reliability.

### Who publishes AI benchmark scores?

Some come from the labs themselves, which is why independent runs matter. Epoch AI publishes its own evaluations under CC-BY, and the Aider polyglot leaderboard is published under Apache 2.0. New models are scored as they ship, so any ranking is a snapshot. [(Epoch AI)](https://epoch.ai/)

### Can AI labs game benchmark results?

Mostly by accident. Benchmark questions leak into training data over time, which inflates scores without improving the model. Labs also choose which benchmarks to report and which prompting setup to use. Independent runs and held-out test sets are the usual defence.

## Next in the learning path

- [Benchmarks The eval suites behind these scores](https://agentfieldbook.org/software/benchmarks/)

- [Models The models being ranked](https://agentfieldbook.org/software/models/)

- [Frameworks & Tools What the top models are built into](https://agentfieldbook.org/software/frameworks/)

- [Safety Evals Scoring safety, not only capability](https://agentfieldbook.org/security/evals/)

- [Harnesses The SDKs that run these models](https://agentfieldbook.org/software/frameworks/)

## Data

- [benchmark_scores.json](https://agentfieldbook.org/data/benchmark_scores.json)

Licensed [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/).
