Corpus Agentis
The field book to agent ecosystems
The field book to agent ecosystems
Compute · Chips

AI accelerators by FP16 performance

The silicon agents ultimately run on, top accelerators by FP16 throughput (Epoch AI, CC-BY). Bandwidth & power in the full dataset.

AMD Instinct MI355X · AMD2.52 PFLOP/s
NVIDIA GB300 (Blackwell Ultra) · NVIDIA2.50 PFLOP/s
NVIDIA GB200 · NVIDIA2.50 PFLOP/s
Google TPU v7 Ironwood · Google2.31 PFLOP/s
AMD Instinct MI350X · AMD2.31 PFLOP/s
NVIDIA B300 (Blackwell Ultra) · NVIDIA2.25 PFLOP/s
NVIDIA B200 · NVIDIA2.25 PFLOP/s
Intel Habana Gaudi3 · Intel1.68 PFLOP/s
AMD Instinct MI325X · AMD1.31 PFLOP/s
AMD Instinct MI300X · AMD1.31 PFLOP/s
Microsoft Maia 200 · Microsoft1.27 PFLOP/s
Biren BR100 · Biren1.02 PFLOP/s
NVIDIA H200 SXM · NVIDIA990 TFLOP/s
NVIDIA GH200 · NVIDIA990 TFLOP/s
NVIDIA GH100 · NVIDIA990 TFLOP/s
NVIDIA H800 SXM5 · NVIDIA989 TFLOP/s
NVIDIA H100 SXM5 80GB · NVIDIA989 TFLOP/s
AMD Instinct MI300A · AMD981 TFLOP/s
Google TPU v6e Trillium · Google918 TFLOP/s
Huawei Ascend 920 · Huawei900 TFLOP/s
Chip designers in the network · in the supply-chain network
NVIDIA (H100/B200) · ~87% AI trainingAMD (MI300X)Google TPUAWS Trainium2Cerebras WSE-3 · 450+ tok/sGroq LPU · 276-330 tok/sHuawei Ascend · capped at SMIC 7nmMicrosoft Maia · Maia 100, Azure custom AI chip (TSMC N5, 64GB HBM2E)AWS Inferentia · Inferentia2, AWS inference chip (EC2 Inf2)Intel Gaudi · Gaudi 3, Intel AI accelerator (TSMC 5nm)SambaNova SN40L · Reconfigurable Dataflow Unit (TSMC 5nm)Tenstorrent Blackhole · RISC-V AI accelerator (TSMC 6nm)Meta MTIA · Meta Training & Inference Accelerator (Broadcom co-design)
Common questions
Why are AI chips so expensive?

Because supply is tight and they are hard to make. Frontier accelerators are built on the newest process node at a handful of factories, each one carries scarce high-bandwidth memory, and the packaging that bonds the two together is itself a bottleneck. Demand from every cloud provider competes for the same output.

What is the difference between a GPU and a TPU?

A GPU is a general-purpose parallel processor that was adapted for AI. A TPU is a chip built only for the tensor maths that neural networks use. In practice the choice comes down to throughput, on-chip memory, and whether your model and serving stack run on that silicon at all.

What is FP16 and why is it used to compare AI chips?

FP16 is half-precision floating point, a 16-bit number format. AI models mostly run in this kind of reduced precision, so FP16 operations per second tracks real workload speed better than a headline peak-FLOP figure does. The figures ranked here come from an open accelerator dataset (Epoch AI).

Does Nvidia make its own chips?

No. Nvidia designs its accelerators and TSMC manufactures them. Almost every AI chip company is fabless, meaning it owns the design but not the factory, which is why foundry capacity limits how many chips can exist.

How much memory does an AI chip need?

Enough to hold the model, or it has to shuffle data in and out and slows right down. Frontier accelerators ship with tens to hundreds of gigabytes of high-bandwidth memory on the package. Bandwidth matters as much as capacity, because the compute units sit idle waiting for memory to feed them.

Next in the learning path