Real benchmark results across the capabilities agents are hired for: coding, science/reasoning, math, factuality. Epoch AI (CC-BY) + Aider Polyglot (Apache-2.0), live. Pick a capability.
Aider Polyglot CodingGPQA Diamond Science / ReasoningAIME MathFrontierMath Frontier MathSimpleQA FactualitySWE-bench Verified CodingMMMU MultimodalOSWorld Computer UseArtificial Analysis Image Arena ImageArtificial Analysis Video Arena VideoArtificial Analysis Speech Arena AudioArtificial Analysis Output Speed Speed
Aider Polyglot, Multi-language code editing across a polyglot repository.
Coding · Aider Polyglot Live data
4.gemini-2.5-pro-preview-06-05 (32k think)83.1%
8.gemini-2.5-pro-preview-06-05 (default think)79.1%
9.o3 (high) + gpt-4.178.2%
10.Gemini 2.5 Pro Preview 05-0676.9%
12.DeepSeek-V3.2-Exp (Reasoner)74.2%
13.Gemini 2.5 Pro Preview 03-2572.9%
15.claude-opus-4-20250514 (32k thinking)72%
Bar baseline zoomed to 65.9% so ranking reads · numbers are the true values
GPQA Diamond, Graduate-level science Q&A (biology, physics, chemistry), "Google-proof" questions experts get right and laypeople can't.
Science / Reasoning · GPQA Diamond Live data
1.gpt-5.4-pro-2026-03-05_xhigh94.6%
2.gemini-3.1-pro-preview94.1%
3.gpt-5.5-pre-release_xhigh94%
4.gpt-5.5-pro-pre-release_xhigh93.9%
5.gpt-5.4-2026-03-05_xhigh93.3%
6.gemini-3.5-flash_high92.8%
7.gemini-3-pro-preview92.6%
9.gpt-5.2-2025-12-11_xhigh91.4%
10.claude-opus-4-8_max91%
13.claude-opus-4-6_32K90.5%
14.claude-opus-4-7_xhigh90.2%
Bar baseline zoomed to 87.6% so ranking reads · numbers are the true values
AIME, Competition mathematics, the American Invitational Mathematics Examination.
Math · AIME Live data
1.gpt-5.5-pro-pre-release_xhigh100%
2.gpt-5.5-pre-release_xhigh100%
3.claude-fable-5_max99.7%
4.claude-opus-4-8_max98.3%
5.claude-opus-4-7_xhigh97.8%
8.gpt-5.2-2025-12-11_high96.1%
9.gpt-5.2-2025-12-11_xhigh96.1%
10.gemini-3.1-pro-preview95.6%
11.gemini-3.5-flash_high95.6%
12.gpt-5.4-2026-03-05_xhigh95.3%
14.claude-opus-4-6_64K94.4%
15.gpt-5.2-2025-12-11_medium93.9%
Bar baseline zoomed to 91.3% so ranking reads · numbers are the true values
FrontierMath, Research-level mathematics, the hardest math benchmark, problems that take specialists hours.
Frontier Math · FrontierMath Live data
3.claude-opus-4-7_xhigh90%
6.gemini-3.1-pro-preview88.9%
7.gemini-3.5-flash_high80%
9.gpt-5.4-2026-03-05_high80%
10.claude-sonnet-4-6_16K80%
11.claude-opus-4-6_32K80%
12.gemini-3-pro-preview80%
13.gpt-5-2025-08-07_medium70%
14.gpt-5.4-nano-2026-03-17_high60%
15.gemini-3-flash-preview60%
Bar baseline zoomed to 45.5% so ranking reads · numbers are the true values
SimpleQA, Short-fact accuracy, does the model hallucinate on simple factual questions?
Factuality · SimpleQA Live data
1.gemini-3.1-pro-preview77.3%
2.gemini-3-pro-preview72.9%
3.gemini-3.5-flash_high68.4%
4.claude-fable-5_xhigh68.3%
5.qwen3-max-2025-09-2367.5%
6.gemini-3-flash-preview67.4%
8.gpt-5.5-pro-pre-release_xhigh64.5%
9.gpt-5.5-pre-release_xhigh63.1%
11.qwen3.6-max-preview56.9%
14.claude-opus-4-7_xhigh50.6%
15.gpt-5-2025-08-07_high50.6%
Bar baseline zoomed to 40.8% so ranking reads · numbers are the true values
SWE-bench Verified, Can it fix real GitHub issues? A human-verified subset of resolved pull requests, the standard for agentic coding.
Coding · SWE-bench Verified Live data
1.claude-opus-4-7_max83.5%
2.gpt-5.5-pre-release_xhigh80.6%
3.gemini-3.5-flash_high79.3%
5.gpt-5.4-2026-03-05_high76.9%
6.qwen3.6-max-preview76.7%
8.claude-opus-4-5-2025110176.7%
9.gemini-3.1-pro-preview-customtools75.6%
10.gemini-3-flash-preview75.4%
11.claude-sonnet-4-675.2%
12.gpt-5.3-codex_high74.8%
15.gpt-5.2-2025-12-11_high73.8%
Bar baseline zoomed to 69.9% so ranking reads · numbers are the true values
MMMU, Multimodal college-level reasoning across images, diagrams and text.
Multimodal · MMMU Live data
1.Gemini 3.1 Pro Preview88.21%
5.Claude Sonnet 4.683.58%
6.Claude Opus 4.6 (Thinking)83.18%
Bar baseline zoomed to 78.4% so ranking reads · numbers are the true values
OSWorld, Real computer-use tasks in a live operating system, files, browser, apps.
Computer Use · OSWorld Live data
1.Specialist model (SOTA)80.4%
Bar baseline zoomed to 60.9% so ranking reads · numbers are the true values
Image · Artificial Analysis Image Arena Live data
1.GPT Image 2 (high)1339%
4.HiDream-O1-Image-1.51265%
5.Nano Banana 2 Lite1262%
6.Cosmos3-Super-Text2Image1226%
7.HiDream-O1-Image-Dev-26041186%
8.Ideogram 4.0 Quality1168%
Bar baseline zoomed to 1107.7% so ranking reads · numbers are the true values
Video · Artificial Analysis Video Arena Live data
1.Dreamina Seedance 2.0 720p1224%
3.Kling 3.0 1080p (Pro)1111%
9.grok-imagine-video1072%
Bar baseline zoomed to 909.0% so ranking reads · numbers are the true values
Audio · Artificial Analysis Speech Arena Live data
2.Gemini 3.1 Flash TTS1214%
Bar baseline zoomed to 989.1% so ranking reads · numbers are the true values
Speed · Artificial Analysis Output Speed Live data
2.Granite 4.0 H Small473%
5.Gemini 3.1 Flash-Lite338%
Models behind these scores · in the supply-chain network
OpenAI GPT · ~25% enterprise prod share (Menlo, mid-2025)Anthropic Claude · ~32% enterprise prod share, ahead of OpenAI (Menlo, mid-2025)Google GeminiMeta Llama · OSS leaderMistral · EU altDeepSeek · cheapest frontierxAI Grok · Grok 4.x frontier modelsAlibaba Qwen · Open-weight Qwen3 seriesAmazon Nova · AWS Nova FM family (Bedrock)Cohere · Command enterprise LLMs (RAG / agents)