Corpus Agentis
The field book to agent ecosystems
The field book to agent ecosystems
Compute · Cloud & Infra

Where agents run

The clouds and inference providers hosting agent workloads. The bars give the market structure; the who-runs-where table traces which agents, models and platforms sit on which cloud, from the relationship data. Every provider is sourced: click through for the citation.

hyperscaler · 4
◆ 20

AWS (Bedrock)

Note
Anthropic primary cloud
◆ 18

Microsoft Azure

Note
OpenAI primary cloud
◆ 11

Google Cloud

Note
Gemini + Anthropic (TPU)
◆ 7

Oracle (OCI)

Note
130k+ GPU clusters
Neocloud · 4
◆ 3

CoreWeave

Note
First GB300 NVL72 deployment
◆ 1

Nebius

Note
AI-native cloud with NVIDIA H100 / B200 clusters; NVIDIA reference platform partner.

Lambda

Note
NVIDIA GPU cloud, HGX B200 / H100 for training and inference; popular for OSS agent shops.
◆ 1

Crusoe

Note
Sustainable NVIDIA GPU cloud running on stranded energy: H100 / B200 clusters for training + inference.
Edge / serverless · 5

Cloudflare (Workers AI)

Note
Workers + Durable Objects + Workers AI running agents at the edge. Agents SDK, low cold start, GPU inference in 190+ cities.

Netlify

Note
Edge Functions + Serverless Functions running agent workloads. Native LLM adapters + AI SDK partners.

Vercel

Note
AI SDK, Edge + Serverless functions, streaming responses, function-calling helpers: the default deploy target for Next.js agent apps.

Fly.io

Note
Region-anchored VMs for agent runtimes; popular for LangGraph deployments needing durable state + low latency.

Deno Deploy

Note
Edge JavaScript runtime for TypeScript agents: used for MCP servers, Slack bots, and lightweight tool-call routers.
PaaS · 2

Railway

Note
App-hosting PaaS with git-integrated deploys: common for LangChain + LlamaIndex prototype-to-prod.

Render

Note
Managed app + background-worker hosting used by LangChain + LlamaIndex deploys with git-integrated pipelines.
GPU / sandbox · 1

Modal

Note
GPU-first Python cloud with sandboxed compute: popular for coding-agent workloads + heavy inference.
Inference · 6
◆ 4

Together AI

Note
Model inference across 200+ open-source models: the default OSS-model provider for agents. Backed the MoA paper.
◆ 5

Fireworks AI

Note
Fast inference for open models with function-calling helpers, LoRA fine-tunes, tool-use benchmarks.

GroqCloud

Note
Extreme-low-latency LPU inference, the go-to for latency-critical agent surfaces (voice, real-time tool-use).

Cerebras

Note
Wafer-scale inference, record-holder for fastest Llama / GPT-oss inference; used by latency-critical agents.

HuggingFace Inference Endpoints

Note
Managed inference for any HF model, one-click deploy across AWS/Azure/GCP regions; the default OSS-model endpoint for agents.

Baseten

Note
Model + agent inference platform with Truss packaging; used for coding-agent + voice-agent workloads.
GPU / Ray · 1

Anyscale

Note
Managed Ray for distributed agent workloads: training, batch inference, multi-agent orchestration at scale.
GPU / spot · 2

RunPod

Note
On-demand + spot GPU cloud, H100 / A100 / L40S at competitive prices; popular for OSS agent + fine-tune workloads.

Vast.ai

Note
Community-sourced GPU marketplace, cheapest H100 / RTX 4090 rentals; used for research + spot inference.
GPU · 1

DigitalOcean GPU Droplets

Note
Managed NVIDIA GPU VMs (H100 / L40S) for small-team agent inference + fine-tune workloads.
Market structure · providers by category
Inference
6
Edge / serverless
5
hyperscaler
4
Neocloud
4
PaaS
2
GPU / spot
2
GPU / sandbox
1
GPU / Ray
1
GPU
1
Who runs where · 9 providers with traced workloads
ProviderTypeTraced workloads & dependencies
AWS (Bedrock)hyperscaler20NVIDIA (H100/B200)AWS Trainium2Anthropic ClaudeMarvellAWS InferentiaMeta LlamaMistralDeepSeekCohereAmazon Nova+10 more
Microsoft Azurehyperscaler18NVIDIA (H100/B200)AMD (MI300X)OpenAI GPTMistralBroadcomArista NetworksMicrosoft MaiaMeta LlamaDeepSeekCohere+8 more
Google Cloudhyperscaler11NVIDIA (H100/B200)Google TPUAnthropic ClaudeGoogle GeminiMeta LlamaMistralAlibaba QwenDeepSeekKairos PowerSchneider Electric+1 more
Oracle (OCI)hyperscaler7NVIDIA (H100/B200)AMD (MI300X)CohereMeta LlamaxAI GrokOpenAI GPTGoogle Gemini
Fireworks AIInference5NVIDIA (H100/B200)Meta LlamaDeepSeekAlibaba QwenMistral
Together AIInference4NVIDIA (H100/B200)Meta LlamaDeepSeekAlibaba Qwen
CoreWeaveNeocloud3NVIDIA (H100/B200)OpenAI GPTVertiv
NebiusNeocloud1NVIDIA (H100/B200)
CrusoeNeocloud1NVIDIA (H100/B200)
Full dataset · 26 providers
26 rows Live data · click a row for its full spec
namecountrytypesrc
AWS (Bedrock)UShyperscaler◆20↗ T1
Microsoft AzureUShyperscaler◆18↗ T1
Google CloudUShyperscaler◆11↗ T1
CoreWeaveUSNeocloud◆3↗ T1
Oracle (OCI)UShyperscaler◆7↗ T1
Cloudflare (Workers AI)Edge / serverless↗ T1
NetlifyEdge / serverless↗ T1
VercelEdge / serverless↗ T1
Fly.ioEdge / serverless↗ T1
RailwayPaaS↗ T1
ModalGPU / sandbox↗ T1
Together AIInference◆4↗ T1
Fireworks AIInference◆5↗ T1
GroqCloudInference↗ T1
AnyscaleGPU / Ray↗ T1
NebiusNeocloud◆1↗ T1
LambdaNeocloud↗ T1
CrusoeNeocloud◆1↗ T1
CerebrasInference↗ T1
HuggingFace Inference EndpointsInference↗ T1
BasetenInference↗ T1
RunPodGPU / spot↗ T1
DigitalOcean GPU DropletsGPU↗ T1
Vast.aiGPU / spot↗ T1
RenderPaaS↗ T1
Deno DeployEdge / serverless↗ T1
Clouds & inference in the network · in the supply-chain network
AWS (Bedrock)Microsoft AzureGoogle CloudCoreWeaveTogether AIFireworks AIGroqCloudOracle (OCI) · 130k+ GPU clustersHuawei Cloud · CloudMatrix (Ascend 910C)Lambda · NVIDIA GPU cloud (HGX B200/H100)Nebius · AI cloud, NVIDIA Reference Platform partnerCrusoe · GPU cloud (GB200/B200/H200)SambaNova Cloud · SambaNova hosted RDU cloud
Field notes

Hyperscalers, neoclouds and inference hosts, and which workloads sit on each

Where an agent runs is one of the more consequential decisions in the build, and the answer splits by what the agent is for. The frontier models themselves sit on a short list of hyperscalers, AWS, Microsoft Azure, Google Cloud and Oracle, and for internal software, applied AI inside a SaaS product or anything already sitting next to corporate data, that is usually the right place to put the workload too.

Show more

Customer-facing agents pull the other way. A quick FAQ responder or anything judged on how fast it answers is better served at the edge, where Cloudflare, Vercel, Netlify and Fly.io compete on cold start and proximity rather than on GPU-hour price. Neoclouds sit in between, cheaper per GPU-hour for sustained training and serving. The pattern across all of it is workload splitting rather than consolidation: training, serving and the agent loop increasingly land in different places.

From the corpus, curated by Brandon Chaplin
Common questions
Where do AI agents actually run?

On cloud infrastructure rented from a provider. In practice that means one of four kinds: hyperscalers such as AWS, Azure and Google Cloud, specialist GPU providers, edge and serverless platforms close to the user, or a managed inference host. Which one fits depends on latency, where the data must sit, and cost.

What is a neocloud?

A neocloud is a specialist provider built around GPU capacity for AI, rather than a general-purpose cloud. CoreWeave, Lambda and Nebius are examples. They often get new silicon earlier and price aggressively per GPU-hour, but offer far fewer managed services.

How much does it cost to rent a GPU in the cloud?

Prices are quoted per GPU-hour and vary widely by chip and provider. Top-end accelerators run to several dollars an hour on demand, and less on committed or spot contracts. Specialist providers usually undercut the hyperscalers on raw compute.

Is it cheaper to rent a GPU or pay per token?

Per token, for most workloads. Renting a machine means paying for it whether or not it is busy, and agent traffic is bursty. Dedicated capacity only wins at high, steady volume, or when you need to run a model nobody hosts for you.

Does it matter which cloud region an agent runs in?

Yes, for two reasons. Distance adds latency to every step, and an agent loop pays that delay on each call rather than once. Data protection rules may also require that personal data is processed in a particular country.

Next in the learning path