Open SourceAPI Host BenchmarksEstimated
Quasar 438B — API Providers & Speed Matrix
Modelled throughput, TTFT latency and prompt-cache pricing estimates across cloud hosting endpoints, derived from official list prices and public benchmarks.
Fastest provider
1188 t/s
Lowest latency (TTFT)
90 ms P50
Best price
$0.18 / $0.8
Blended Pricing Workload Calculator
Switch between standard conversational ratios, agentic caching loops, and heavy RAG workloads.
| Provider & Host | Hardware Stack | Output Speed (TPS) | Latency TTFT (P50) | Input Price ($/1M) | Output Price ($/1M) | Cache Read ($/1M) | Blended Price ($/1M) | SLA |
|---|---|---|---|---|---|---|---|---|
Open SourceOfficial | Cloud TPU / Nvidia H100 | 95 t/s | 800 ms (1.12 s p90) | $0.4 | $1.6 | $0.04 | $0.7 | 99.99% |
| Groq LPU v2 (SRAM) | 399 t/s | 90 ms (120 ms p90) | $0.34 | $1.36 | $0.032 | $0.59 | 99.95% | |
| CS-3 Wafer-Scale Engine | 1188 t/s | 120 ms (160 ms p90) | $0.36 | $1.44 | — | $0.63 | 99.90% | |
| Nvidia H100 SXM5 | 143 t/s | 640 ms (880 ms p90) | $0.26 | $1.12 | $0.024 | $0.48 | 99.95% | |
| Nvidia H100 / L40S | 124 t/s | 720 ms (1.00 s p90) | $0.18 | $0.8 | $0.02 | $0.34 | 99.90% | |
| Nvidia B200 / H100 | 181 t/s | 560 ms (760 ms p90) | $0.24 | $1.04 | $0.024 | $0.44 | 99.99% |
Figures are modelled estimates based on published list prices and public throughput benchmarks — not measured on live endpoints. Verify current pricing and terms with each host before committing.