---
title: "AI inference throughput · tokens per second and model-access platforms · Agent Fieldbook"
url: https://agentfieldbook.org/hardware/inference/
description: "Independent inference throughput (Artificial Analysis), tagged with the model measured, plus the model-access platforms and edge inference catalog agents draw models from. A source-graded field book to how fast agents run and where they reach a model."
section: "Compute · Inference"
source: Agent Fieldbook — generated from the published page
---

# Inference throughput

**Managed cloud catalogue · Edge-hosted · Gateway (proxy) · Model-choice layer · deepseek-r1-distill-qwen-32b · gemma-2b-it-lora · gemma-3-12b-it · gemma-4-26b-a4b-it ƒ · gemma-7b-it-lora · gemma-sea-lion-v4-27b-it · glm-4.7-flash ƒ · gpt-oss-120b ƒ · gpt-oss-20b ƒ · granite-4.0-h-micro ƒ · kimi-k2.5 ƒ · kimi-k2.6 ƒ · kimi-k2.7-code ƒ · llama-2-7b-chat-fp16 · llama-2-7b-chat-hf-lora · llama-2-7b-chat-int8 · llama-3-8b-instruct · llama-3-8b-instruct-awq · llama-3.1-70b-instruct · llama-3.1-8b-instruct · llama-3.1-8b-instruct-awq · llama-3.1-8b-instruct-fast · llama-3.1-8b-instruct-fp8 · llama-3.2-11b-vision-instruct · llama-3.2-1b-instruct · llama-3.2-3b-instruct · llama-3.3-70b-instruct-fp8-fast ƒ · llama-4-scout-17b-16e-instruct ƒ · llama-guard-3-8b · mistral-7b-instruct-v0.1 · mistral-7b-instruct-v0.2-lora · mistral-small-3.1-24b-instruct ƒ · nemotron-3-120b-a12b ƒ · phi-2 · qwen2.5-coder-32b-instruct · qwen3-30b-a3b-fp8 ƒ · qwq-32b · sqlcoder-7b-2 · bge-base-en-v1.5 · bge-large-en-v1.5 · bge-m3 · bge-small-en-v1.5 · embeddinggemma-300m · plamo-embedding-1b · qwen3-embedding-0.6b · flux · nova-3 · whisper · whisper-large-v3-turbo · whisper-tiny-en ƒ · aura-1 · aura-2-en · aura-2-es · melotts · dreamshaper-8-lcm · flux-1-schnell · flux-2-dev · flux-2-klein-4b · flux-2-klein-9b · lucid-origin · phoenix-1.0 · stable-diffusion-v1-5-img2img · stable-diffusion-v1-5-inpainting · stable-diffusion-xl-base-1.0 · stable-diffusion-xl-lightning · llava-1.5-7b-hf · uform-gen2-qwen-500m · indictrans2-en-indic-1B · m2m100-1.2b · bart-large-cnn · bge-reranker-base · distilbert-sst-2-int8 · detr-resnet-50 · resnet-50 · smart-turn-v2 · AWS (Bedrock) · Microsoft Azure · Google Cloud · CoreWeave · Together AI · Fireworks AI · GroqCloud · Oracle (OCI) · 130k+ GPU clusters · Huawei Cloud · CloudMatrix (Ascend 910C) · Lambda · NVIDIA GPU cloud (HGX B200/H100) · Nebius · AI cloud, NVIDIA Reference Platform partner · Crusoe · GPU cloud (GB200/B200/H200) · SambaNova Cloud · SambaNova hosted RDU cloud**

Independent output speed (Artificial Analysis), tagged with the model measured, tokens/sec swings 5–20x by model size, so a single unqualified figure misleads. Bars show the headline number; the table carries the sourcing.

## 9 rows Live data · click a row for its full spec

| name | country | type | throughput | src |
| --- | --- | --- | --- | --- |
| Groq | US | inference_api | ~321 t/s· Llama 3.3 70B (≈800 t/s on 8B-class) | [↗ T1](https://artificialanalysis.ai/models/llama-3-3-instruct-70b/providers)[↗Groq](https://groq.com) |
| Cerebras | US | inference_api | ~2,314 t/s· Llama 3.3 70B (~981 t/s on Kimi K2.6) | [↗ T1](https://artificialanalysis.ai/models/kimi-k2-6/providers)[↗Cerebra](https://cerebras.ai) |
| Together AI | US | inference_api | ~198 t/s· Kimi K2.6 (35–434 t/s across models) | ◆4 [↗ T1](https://artificialanalysis.ai/models/kimi-k2-6/providers)[↗Togethe](https://together.ai) |
| Fireworks AI | US | inference_api | ~373 t/s· Kimi K2.6 (up to 740 t/s on gpt-oss-120b) | ◆5 [↗ T1](https://artificialanalysis.ai/models/kimi-k2-6/providers)[↗Firewor](https://fireworks.ai) |
| SambaNova | US | inference_api | ~691 t/s· gpt-oss-120b (high) | ◆2 [↗ T1](https://artificialanalysis.ai/models/gpt-oss-120b/providers)[↗SambaNo](https://sambanova.ai) |
| Nebius | NL | inference_api | ~343 t/s· gpt-oss-120b (high) | ◆1 [↗ T1](https://artificialanalysis.ai/models/gpt-oss-120b/providers)[↗Nebius](https://nebius.com) |
| Baseten | US | inference_api | ~223 t/s· gpt-oss-120b (high) | [↗ T1](https://artificialanalysis.ai/models/gpt-oss-120b/providers)[↗Baseten](https://baseten.co) |
| DeepInfra | US | inference_api | ~154 t/s· gpt-oss-120b (high, Turbo) | [↗ T1](https://artificialanalysis.ai/models/gpt-oss-120b/providers)[↗DeepInf](https://deepinfra.com) |
| Novita AI | US | inference_api | ~133 t/s· gpt-oss-120b (high) | [↗ T1](https://artificialanalysis.ai/models/gpt-oss-120b/providers)[↗Novita](https://novita.ai) |

## Model-access platforms · 8 · where agents get their models

Beyond a single host: the managed cloud catalogues, edge hosts and gateways an agent draws models from. Cloud catalogues serve their own line-up; gateways proxy hundreds behind one endpoint. Every count is from the platform's own docs.

| Platform | Kind | Models | Notable | src |
| --- | --- | --- | --- | --- |
| **Amazon Bedrock** Amazon Web Services | Managed cloud catalogue | 100+ models | Anthropic Claude· Meta Llama· Mistral· Cohere· AI21· Amazon Nova/Titan· DeepSeek· OpenAI· Stability | [↗ T1](https://docs.aws.amazon.com/bedrock/latest/userguide/models-supported.html) |
| **Azure AI Foundry (Microsoft Foundry)** Microsoft | Managed cloud catalogue | 1,900+ sold directly (11K+ incl. community) | Azure OpenAI· Claude· Meta· Mistral· DeepSeek· xAI Grok· Cohere· NVIDIA· Fireworks· Hugging Face | [↗ T1](https://learn.microsoft.com/en-us/azure/ai-foundry/concepts/foundry-models-overview) |
| **Google Vertex AI Model Garden** Google | Managed cloud catalogue | 200+ models | Gemini· Claude· Llama· Mistral· Gemma· plus curated open models | [↗ T1](https://cloud.google.com/vertex-ai/generative-ai/docs/model-garden/available-models) |
| **Cloudflare Workers AI** Cloudflare | Edge-hosted | ~50 open models | Llama· Mistral· Qwen· Gemma· embeddings + vision, served at the edge | [↗ T1](https://developers.cloudflare.com/workers-ai/models/) |
| **Netlify AI Gateway** Netlify | Gateway (proxy) | Multi-provider | Provider proxy with usage controls for agents on Netlify Functions + Edge | [↗ T2](https://www.netlify.com/platform/ai-gateway/) |
| **OpenRouter** OpenRouter | Gateway (proxy) | 300+ models | One API key across every major provider· routing, load-balancing, fallbacks | [↗ T1](https://openrouter.ai/models) |
| **Vercel AI Gateway** Vercel | Gateway (proxy) | Hundreds of models | Unified API across providers· text/image/video/embeddings· BYOK, budgets, fallbacks | [↗ T1](https://vercel.com/ai-gateway/models) |
| **Salesforce Einstein / Agentforce** Salesforce | Model-choice layer | Curated + bring-your-own-LLM | Model choice behind the Einstein Trust Layer (masking, zero-retention)· Models API· BYOLLM | [↗ T2](https://www.salesforce.com/products/platform/einstein-trust-layer/) |

## Edge inference catalog · 75 models on Cloudflare Workers AI Live data

The reasoning cores and **capability models**, voice, vision, embeddings, that agents plug into at the edge. A supporting layer: agents are the subject; these are what powers them.

## Text Generation · 38

## Text Embeddings · 7

## Automatic Speech Recognition · 5

## Text-to-Speech · 4

## Text-to-Image · 11

## Image-to-Text · 2

## Translation · 2

## Summarization · 1

## Text Classification · 2

## Object Detection · 1

## Image Classification · 1

## Voice Activity Detection · 1

## Clouds & inference in the network · in the supply-chain network

## Field notes

Independent tokens-per-second, plus where agents get their models

Nine inference providers sit between the model and the application, selling tokens rather than machines. For most teams this is the layer they actually buy.

Competition here is on latency and throughput more than price, because the price floor is set by hardware everyone rents from the same suppliers. Time-to-first-token is the number that matters for agents, since a loop that calls the model ten times pays that latency ten times.

- [Brandon Chaplin](https://www.linkedin.com/in/brandon-chaplin-digital-marketing-strategist)

## Common questions

### What is inference in AI?

Inference is running a trained model to get an answer, as opposed to training it in the first place. Every time an agent calls a model, that is one inference. It is the cost a production system actually pays, over and over.

### What is a token, and what is tokens per second?

A token is a chunk of text, roughly three-quarters of a word on average. Tokens per second is how fast a model produces its answer. Independent benchmarks publish measured speeds per model and provider [(Artificial Analysis)](https://artificialanalysis.ai/).

### What is time to first token?

The delay between sending a request and the first word coming back, as distinct from how fast the rest streams. It matters more for agents than raw speed, because a loop that calls the model ten times waits ten times.

### Why is the same model faster on one provider than another?

Different hardware, different batching, and different quantisation. Providers trade a little accuracy or a lot of cost against speed, so an identical model name can behave very differently depending on who is serving it.

### What is edge inference?

Running models on servers close to the user instead of in one distant region. It cuts round-trip latency and can keep data inside a country. It usually means smaller models, since the biggest ones will not fit at the edge.

## Next in the learning path

- [Cloud & Infra The providers hosting these inference workloads](https://agentfieldbook.org/hardware/cloud_infra/)

- [Chips The accelerators behind the throughput](https://agentfieldbook.org/hardware/chips/)

- [HBM The memory bandwidth that sets the ceiling](https://agentfieldbook.org/hardware/hbm/)

- [Power The energy cost under each token served](https://agentfieldbook.org/hardware/power/)

- [Foundries Where the serving silicon is fabricated](https://agentfieldbook.org/hardware/foundries/)

## Data

- [inference.json](https://agentfieldbook.org/data/inference.json)
- [model_platforms.json](https://agentfieldbook.org/data/model_platforms.json)
- [inference_catalog.json](https://agentfieldbook.org/data/inference_catalog.json)

Licensed [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/).
