AI Hardware & Silicon Inference Leaderboard

Standardized benchmarks comparing Wafer-Scale Engines, Tensor Streaming LPUs, RDUs, and GPU clusters on Llama 3.3 & DeepSeek-R1 inference throughput.

Hardware PlatformSilicon ArchitecturePeak Memory BandwidthPeak Output Speed (TPS)TTFT Prefill (2k context)Scalability DomainStatus
Cerebras CS-3
Cerebras Systems
Wafer-Scale Engine (WSE-3)9,000 TB/s (9 PB/s on-chip SRAM)1850 t/s0.12sLinear wafer clustersIn Production
Groq LPU v2
Groq
Tensor Streaming Processor (TSP)80 TB/s (Deterministic SRAM)520 t/s0.08sRack-scale meshIn Production
Nvidia GB200 NVL72
Nvidia
Blackwell GPU + Grace CPU576 TB/s aggregate HBM3e380 t/s0.16sNVLink 5 domain (72 GPUs)In Production
SambaNova SN40L
SambaNova Systems
Reconfigurable Dataflow Unit (RDU)3-tier (SRAM + HBM + DDR5)310 t/s0.19sDataScale SN40L NodeIn Production
Google TPU v5p
Google Cloud
Tensor Processing Unit (v5p OCS)4.8 TB/s HBM3 per chip240 t/s0.22s8,960 chip Pod via OCSIn Production
Nvidia H100 SXM5
Nvidia
Hopper GPU Architecture3.35 TB/s HBM3160 t/s0.26s8-GPU HGX clusterIn Production
0 / 4