Semiconductor & Inference Hardware
AI Hardware & Silicon Inference Leaderboard
Standardized benchmarks comparing Wafer-Scale Engines, Tensor Streaming LPUs, RDUs, and GPU clusters on Llama 3.3 & DeepSeek-R1 inference throughput.
| Hardware Platform | Silicon Architecture | Peak Memory Bandwidth | Peak Output Speed (TPS) | TTFT Prefill (2k context) | Scalability Domain | Status |
|---|---|---|---|---|---|---|
Cerebras CS-3 Cerebras Systems | Wafer-Scale Engine (WSE-3) | 9,000 TB/s (9 PB/s on-chip SRAM) | 1850 t/s | 0.12s | Linear wafer clusters | In Production |
Groq LPU v2 Groq | Tensor Streaming Processor (TSP) | 80 TB/s (Deterministic SRAM) | 520 t/s | 0.08s | Rack-scale mesh | In Production |
Nvidia GB200 NVL72 Nvidia | Blackwell GPU + Grace CPU | 576 TB/s aggregate HBM3e | 380 t/s | 0.16s | NVLink 5 domain (72 GPUs) | In Production |
SambaNova SN40L SambaNova Systems | Reconfigurable Dataflow Unit (RDU) | 3-tier (SRAM + HBM + DDR5) | 310 t/s | 0.19s | DataScale SN40L Node | In Production |
Google TPU v5p Google Cloud | Tensor Processing Unit (v5p OCS) | 4.8 TB/s HBM3 per chip | 240 t/s | 0.22s | 8,960 chip Pod via OCS | In Production |
Nvidia H100 SXM5 Nvidia | Hopper GPU Architecture | 3.35 TB/s HBM3 | 160 t/s | 0.26s | 8-GPU HGX cluster | In Production |