Corpus Agentis
The field book to agent ecosystems
The field book to agent ecosystems
Compute · Inference

Inference throughput

Independent output speed (Artificial Analysis), tagged with the model measured, tokens/sec swings 5–20x by model size, so a single unqualified figure misleads. Bars show the headline number; the table carries the sourcing.

SambaNova~691 t/s · gpt-oss-120b (high)
Nebius~343 t/s · gpt-oss-120b (high)
Baseten~223 t/s · gpt-oss-120b (high)
DeepInfra~154 t/s · gpt-oss-120b (high, Turbo)
Novita AI~133 t/s · gpt-oss-120b (high)
Cerebras~2,314 t/s · Llama 3.3 70B (~981 t/s on Kimi K2.6)
Fireworks AI~373 t/s · Kimi K2.6 (up to 740 t/s on gpt-oss-120b)
Groq~321 t/s · Llama 3.3 70B (≈800 t/s on 8B-class)
Together AI~198 t/s · Kimi K2.6 (35–434 t/s across models)
9 rows Live data · click a row for its full spec
namecountrytypethroughputsrc
GroqUSinference_api~321 t/s · Llama 3.3 70B (≈800 t/s on 8B-class)↗ T1↗Groq
CerebrasUSinference_api~2,314 t/s · Llama 3.3 70B (~981 t/s on Kimi K2.6)↗ T1↗Cerebra
Together AIUSinference_api~198 t/s · Kimi K2.6 (35–434 t/s across models)◆4↗ T1↗Togethe
Fireworks AIUSinference_api~373 t/s · Kimi K2.6 (up to 740 t/s on gpt-oss-120b)◆5↗ T1↗Firewor
SambaNovaUSinference_api~691 t/s · gpt-oss-120b (high)◆2↗ T1↗SambaNo
NebiusNLinference_api~343 t/s · gpt-oss-120b (high)◆1↗ T1↗Nebius
BasetenUSinference_api~223 t/s · gpt-oss-120b (high)↗ T1↗Baseten
DeepInfraUSinference_api~154 t/s · gpt-oss-120b (high, Turbo)↗ T1↗DeepInf
Novita AIUSinference_api~133 t/s · gpt-oss-120b (high)↗ T1↗Novita
Model-access platforms · 8 · where agents get their models

Beyond a single host: the managed cloud catalogues, edge hosts and gateways an agent draws models from. Cloud catalogues serve their own line-up; gateways proxy hundreds behind one endpoint. Every count is from the platform's own docs.

PlatformKindModelsNotablesrc
Amazon Bedrock
Amazon Web Services
Managed cloud catalogue100+ modelsAnthropic Claude · Meta Llama · Mistral · Cohere · AI21 · Amazon Nova/Titan · DeepSeek · OpenAI · Stability↗ T1
Azure AI Foundry (Microsoft Foundry)
Microsoft
Managed cloud catalogue1,900+ sold directly (11K+ incl. community)Azure OpenAI · Claude · Meta · Mistral · DeepSeek · xAI Grok · Cohere · NVIDIA · Fireworks · Hugging Face↗ T1
Google Vertex AI Model Garden
Google
Managed cloud catalogue200+ modelsGemini · Claude · Llama · Mistral · Gemma · plus curated open models↗ T1
Cloudflare Workers AI
Cloudflare
Edge-hosted~50 open modelsLlama · Mistral · Qwen · Gemma · embeddings + vision, served at the edge↗ T1
Netlify AI Gateway
Netlify
Gateway (proxy)Multi-providerProvider proxy with usage controls for agents on Netlify Functions + Edge↗ T2
OpenRouter
OpenRouter
Gateway (proxy)300+ modelsOne API key across every major provider · routing, load-balancing, fallbacks↗ T1
Vercel AI Gateway
Vercel
Gateway (proxy)Hundreds of modelsUnified API across providers · text/image/video/embeddings · BYOK, budgets, fallbacks↗ T1
Salesforce Einstein / Agentforce
Salesforce
Model-choice layerCurated + bring-your-own-LLMModel choice behind the Einstein Trust Layer (masking, zero-retention) · Models API · BYOLLM↗ T2
Edge inference catalog · 75 models on Cloudflare Workers AI Live data

The reasoning cores and capability models, voice, vision, embeddings, that agents plug into at the edge. A supporting layer: agents are the subject; these are what powers them.

Text Generation · 38
deepseek-r1-distill-qwen-32bgemma-2b-it-loragemma-3-12b-itgemma-4-26b-a4b-it ƒgemma-7b-it-loragemma-sea-lion-v4-27b-itglm-4.7-flash ƒgpt-oss-120b ƒgpt-oss-20b ƒgranite-4.0-h-micro ƒkimi-k2.5 ƒkimi-k2.6 ƒkimi-k2.7-code ƒllama-2-7b-chat-fp16llama-2-7b-chat-hf-lorallama-2-7b-chat-int8llama-3-8b-instructllama-3-8b-instruct-awqllama-3.1-70b-instructllama-3.1-8b-instructllama-3.1-8b-instruct-awqllama-3.1-8b-instruct-fastllama-3.1-8b-instruct-fp8llama-3.2-11b-vision-instructllama-3.2-1b-instructllama-3.2-3b-instructllama-3.3-70b-instruct-fp8-fast ƒllama-4-scout-17b-16e-instruct ƒllama-guard-3-8bmistral-7b-instruct-v0.1mistral-7b-instruct-v0.2-loramistral-small-3.1-24b-instruct ƒnemotron-3-120b-a12b ƒphi-2qwen2.5-coder-32b-instructqwen3-30b-a3b-fp8 ƒqwq-32bsqlcoder-7b-2
Text Embeddings · 7
bge-base-en-v1.5bge-large-en-v1.5bge-m3bge-small-en-v1.5embeddinggemma-300mplamo-embedding-1bqwen3-embedding-0.6b
Automatic Speech Recognition · 5
fluxnova-3whisperwhisper-large-v3-turbowhisper-tiny-en ƒ
Text-to-Speech · 4
aura-1aura-2-enaura-2-esmelotts
Text-to-Image · 11
dreamshaper-8-lcmflux-1-schnellflux-2-devflux-2-klein-4bflux-2-klein-9blucid-originphoenix-1.0stable-diffusion-v1-5-img2imgstable-diffusion-v1-5-inpaintingstable-diffusion-xl-base-1.0stable-diffusion-xl-lightning
Image-to-Text · 2
llava-1.5-7b-hfuform-gen2-qwen-500m
Translation · 2
indictrans2-en-indic-1Bm2m100-1.2b
Summarization · 1
bart-large-cnn
Text Classification · 2
bge-reranker-basedistilbert-sst-2-int8
Object Detection · 1
detr-resnet-50
Image Classification · 1
resnet-50
Voice Activity Detection · 1
smart-turn-v2
Clouds & inference in the network · in the supply-chain network
AWS (Bedrock)Microsoft AzureGoogle CloudCoreWeaveTogether AIFireworks AIGroqCloudOracle (OCI) · 130k+ GPU clustersHuawei Cloud · CloudMatrix (Ascend 910C)Lambda · NVIDIA GPU cloud (HGX B200/H100)Nebius · AI cloud, NVIDIA Reference Platform partnerCrusoe · GPU cloud (GB200/B200/H200)SambaNova Cloud · SambaNova hosted RDU cloud
Field notes

Independent tokens-per-second, plus where agents get their models

Nine inference providers sit between the model and the application, selling tokens rather than machines. For most teams this is the layer they actually buy.

Show more

Competition here is on latency and throughput more than price, because the price floor is set by hardware everyone rents from the same suppliers. Time-to-first-token is the number that matters for agents, since a loop that calls the model ten times pays that latency ten times.

From the corpus, curated by Brandon Chaplin
Common questions
What is inference in AI?

Inference is running a trained model to get an answer, as opposed to training it in the first place. Every time an agent calls a model, that is one inference. It is the cost a production system actually pays, over and over.

What is a token, and what is tokens per second?

A token is a chunk of text, roughly three-quarters of a word on average. Tokens per second is how fast a model produces its answer. Independent benchmarks publish measured speeds per model and provider (Artificial Analysis).

What is time to first token?

The delay between sending a request and the first word coming back, as distinct from how fast the rest streams. It matters more for agents than raw speed, because a loop that calls the model ten times waits ten times.

Why is the same model faster on one provider than another?

Different hardware, different batching, and different quantisation. Providers trade a little accuracy or a lot of cost against speed, so an identical model name can behave very differently depending on who is serving it.

What is edge inference?

Running models on servers close to the user instead of in one distant region. It cuts round-trip latency and can keep data inside a country. It usually means smaller models, since the biggest ones will not fit at the edge.

Next in the learning path