---
title: "Cloud and inference infrastructure · where AI agents run · Agent Fieldbook"
url: https://agentfieldbook.org/hardware/cloud_infra/
description: "The clouds and inference providers hosting AI agent workloads, grouped by category with a who-runs-where map traced from the relationship graph. Hyperscalers, neoclouds, edge platforms and inference hosts, each source-graded."
section: "Compute · Cloud & Infra"
source: Agent Fieldbook — generated from the published page
---

# Where agents run

**NVIDIA (H100/B200) · AWS Trainium2 · Anthropic Claude · Marvell · AWS Inferentia · Meta Llama · Mistral · DeepSeek · Cohere · Amazon Nova · AMD (MI300X) · OpenAI GPT · Broadcom · Arista Networks · Microsoft Maia · Google TPU · Google Gemini · Alibaba Qwen · Kairos Power · Schneider Electric · xAI Grok · Vertiv · AWS (Bedrock) · Microsoft Azure · Google Cloud · CoreWeave · Together AI · Fireworks AI · GroqCloud · Oracle (OCI) · 130k+ GPU clusters · Huawei Cloud · CloudMatrix (Ascend 910C) · Lambda · NVIDIA GPU cloud (HGX B200/H100) · Nebius · AI cloud, NVIDIA Reference Platform partner · Crusoe · GPU cloud (GB200/B200/H200) · SambaNova Cloud · SambaNova hosted RDU cloud**

The clouds and inference providers hosting agent workloads. The bars give the market structure; the who-runs-where table traces which agents, models and platforms sit on which cloud, from the relationship data. Every provider is sourced: click through for the citation.

## hyperscaler · 4

- [↗ source · confirmed](https://aws.amazon.com/bedrock/)

- [↗ source · confirmed](https://azure.microsoft.com)

- [↗ source · confirmed](https://cloud.google.com)

- [↗ source · confirmed](https://www.oracle.com/cloud/)

## Neocloud · 4

- [↗ source · confirmed](https://www.coreweave.com)

- [↗ source · confirmed](https://nebius.com/)

- [↗ source · confirmed](https://lambda.ai/service/gpu-cloud)

- [↗ source · confirmed](https://crusoe.ai/)

## Edge / serverless · 5

- [↗ source · confirmed](https://developers.cloudflare.com/agents/)

- [↗ source · confirmed](https://www.netlify.com/products/functions/)

- [↗ source · confirmed](https://vercel.com/ai)

- [↗ source · confirmed](https://fly.io/)

- [↗ source · confirmed](https://deno.com/deploy)

## PaaS · 2

- [↗ source · confirmed](https://railway.com/)

- [↗ source · confirmed](https://render.com/)

## GPU / sandbox · 1

- [↗ source · confirmed](https://modal.com/)

## Inference · 6

- [↗ source · confirmed](https://www.together.ai/)

- [↗ source · confirmed](https://fireworks.ai/)

- [↗ source · confirmed](https://groq.com/)

- [↗ source · confirmed](https://cerebras.ai/)

- [↗ source · confirmed](https://huggingface.co/inference-endpoints/dedicated)

- [↗ source · confirmed](https://www.baseten.co/)

## GPU / Ray · 1

- [↗ source · confirmed](https://www.anyscale.com/)

## GPU / spot · 2

- [↗ source · confirmed](https://www.runpod.io/)

- [↗ source · confirmed](https://vast.ai/)

## GPU · 1

- [↗ source · confirmed](https://www.digitalocean.com/products/gpu-droplets)

## Market structure · providers by category

## Who runs where · 9 providers with traced workloads

| Provider | Type | ◆ | Traced workloads & dependencies |
| --- | --- | --- | --- |
| **AWS (Bedrock)** | hyperscaler | 20 | NVIDIA (H100/B200) AWS Trainium2 Anthropic Claude Marvell AWS Inferentia Meta Llama Mistral DeepSeek Cohere Amazon Nova +10 more |
| **Microsoft Azure** | hyperscaler | 18 | NVIDIA (H100/B200) AMD (MI300X) OpenAI GPT Mistral Broadcom Arista Networks Microsoft Maia Meta Llama DeepSeek Cohere +8 more |
| **Google Cloud** | hyperscaler | 11 | NVIDIA (H100/B200) Google TPU Anthropic Claude Google Gemini Meta Llama Mistral Alibaba Qwen DeepSeek Kairos Power Schneider Electric +1 more |
| **Oracle (OCI)** | hyperscaler | 7 | NVIDIA (H100/B200) AMD (MI300X) Cohere Meta Llama xAI Grok OpenAI GPT Google Gemini |
| **Fireworks AI** | Inference | 5 | NVIDIA (H100/B200) Meta Llama DeepSeek Alibaba Qwen Mistral |
| **Together AI** | Inference | 4 | NVIDIA (H100/B200) Meta Llama DeepSeek Alibaba Qwen |
| **CoreWeave** | Neocloud | 3 | NVIDIA (H100/B200) OpenAI GPT Vertiv |
| **Nebius** | Neocloud | 1 | NVIDIA (H100/B200) |
| **Crusoe** | Neocloud | 1 | NVIDIA (H100/B200) |

## Full dataset · 26 providers

## 26 rows Live data · click a row for its full spec

| name | country | type | src |
| --- | --- | --- | --- |
| AWS (Bedrock) | US | hyperscaler | ◆20 [↗ T1](https://aws.amazon.com/bedrock/) |
| Microsoft Azure | US | hyperscaler | ◆18 [↗ T1](https://azure.microsoft.com) |
| Google Cloud | US | hyperscaler | ◆11 [↗ T1](https://cloud.google.com) |
| CoreWeave | US | Neocloud | ◆3 [↗ T1](https://www.coreweave.com) |
| Oracle (OCI) | US | hyperscaler | ◆7 [↗ T1](https://www.oracle.com/cloud/) |
| Cloudflare (Workers AI) | – | Edge / serverless | [↗ T1](https://developers.cloudflare.com/agents/) |
| Netlify | – | Edge / serverless | [↗ T1](https://www.netlify.com/products/functions/) |
| Vercel | – | Edge / serverless | [↗ T1](https://vercel.com/ai) |
| Fly.io | – | Edge / serverless | [↗ T1](https://fly.io/) |
| Railway | – | PaaS | [↗ T1](https://railway.com/) |
| Modal | – | GPU / sandbox | [↗ T1](https://modal.com/) |
| Together AI | – | Inference | ◆4 [↗ T1](https://www.together.ai/) |
| Fireworks AI | – | Inference | ◆5 [↗ T1](https://fireworks.ai/) |
| GroqCloud | – | Inference | [↗ T1](https://groq.com/) |
| Anyscale | – | GPU / Ray | [↗ T1](https://www.anyscale.com/) |
| Nebius | – | Neocloud | ◆1 [↗ T1](https://nebius.com/) |
| Lambda | – | Neocloud | [↗ T1](https://lambda.ai/service/gpu-cloud) |
| Crusoe | – | Neocloud | ◆1 [↗ T1](https://crusoe.ai/) |
| Cerebras | – | Inference | [↗ T1](https://cerebras.ai/) |
| HuggingFace Inference Endpoints | – | Inference | [↗ T1](https://huggingface.co/inference-endpoints/dedicated) |
| Baseten | – | Inference | [↗ T1](https://www.baseten.co/) |
| RunPod | – | GPU / spot | [↗ T1](https://www.runpod.io/) |
| DigitalOcean GPU Droplets | – | GPU | [↗ T1](https://www.digitalocean.com/products/gpu-droplets) |
| Vast.ai | – | GPU / spot | [↗ T1](https://vast.ai/) |
| Render | – | PaaS | [↗ T1](https://render.com/) |
| Deno Deploy | – | Edge / serverless | [↗ T1](https://deno.com/deploy) |

## Clouds & inference in the network · in the supply-chain network

## Field notes

Hyperscalers, neoclouds and inference hosts, and which workloads sit on each

Where an agent runs is one of the more consequential decisions in the build, and the answer splits by what the agent is for. The frontier models themselves sit on a short list of hyperscalers, AWS, Microsoft Azure, Google Cloud and Oracle, and for internal software, applied AI inside a SaaS product or anything already sitting next to corporate data, that is usually the right place to put the workload too.

Customer-facing agents pull the other way. A quick FAQ responder or anything judged on how fast it answers is better served at the edge, where Cloudflare, Vercel, Netlify and Fly.io compete on cold start and proximity rather than on GPU-hour price. Neoclouds sit in between, cheaper per GPU-hour for sustained training and serving. The pattern across all of it is workload splitting rather than consolidation: training, serving and the agent loop increasingly land in different places.

- [Brandon Chaplin](https://www.linkedin.com/in/brandon-chaplin-digital-marketing-strategist)

## Common questions

### Where do AI agents actually run?

On cloud infrastructure rented from a provider. In practice that means one of four kinds: hyperscalers such as AWS, Azure and Google Cloud, specialist GPU providers, edge and serverless platforms close to the user, or a managed inference host. Which one fits depends on latency, where the data must sit, and cost.

### What is a neocloud?

A neocloud is a specialist provider built around GPU capacity for AI, rather than a general-purpose cloud. CoreWeave, Lambda and Nebius are examples. They often get new silicon earlier and price aggressively per GPU-hour, but offer far fewer managed services.

### How much does it cost to rent a GPU in the cloud?

Prices are quoted per GPU-hour and vary widely by chip and provider. Top-end accelerators run to several dollars an hour on demand, and less on committed or spot contracts. Specialist providers usually undercut the hyperscalers on raw compute.

### Is it cheaper to rent a GPU or pay per token?

Per token, for most workloads. Renting a machine means paying for it whether or not it is busy, and agent traffic is bursty. Dedicated capacity only wins at high, steady volume, or when you need to run a model nobody hosts for you.

### Does it matter which cloud region an agent runs in?

Yes, for two reasons. Distance adds latency to every step, and an agent loop pays that delay on each call rather than once. Data protection rules may also require that personal data is processed in a particular country.

## Next in the learning path

- [Inference Output speed of the providers hosting workloads](https://agentfieldbook.org/hardware/inference/)

- [Chips The accelerators these clouds rack](https://agentfieldbook.org/hardware/chips/)

- [Power The datacenter energy constraint](https://agentfieldbook.org/hardware/power/)

- [HBM The memory feeding the hosted silicon](https://agentfieldbook.org/hardware/hbm/)

- [Foundries Where the racked silicon is fabricated](https://agentfieldbook.org/hardware/foundries/)

## Data

- [cloud_infra.json](https://agentfieldbook.org/data/cloud_infra.json)

Licensed [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/).
