GoogleAPI Host BenchmarksEstimated
Gemini 3.8 Flash (Medium Reasoning) — API Providers & Speed Matrix
Modelled throughput, TTFT latency and prompt-cache pricing estimates across cloud hosting endpoints, derived from official list prices and public benchmarks.
Gemini 3.8 Flash (Medium Reasoning) is a proprietary model served through Google's official API. Third-party hosting is not available — the matrix below covers the official endpoint only. See the Google provider page for details.
Fastest provider
315 t/s
Lowest latency (TTFT)
160 ms P50
Best price
$0.75 / $3.75
Blended Pricing Workload Calculator
Switch between standard conversational ratios, agentic caching loops, and heavy RAG workloads.
| Provider & Host | Hardware Stack | Output Speed (TPS) | Latency TTFT (P50) | Input Price ($/1M) | Output Price ($/1M) | Cache Read ($/1M) | Blended Price ($/1M) | SLA |
|---|---|---|---|---|---|---|---|---|
GoogleOfficial | Official cloud API | 315 t/s | 160 ms (220 ms p90) | $0.75 | $3.75 | $0.075 | $1.5 | 99.99% |
Figures are modelled estimates based on published list prices and public throughput benchmarks — not measured on live endpoints. Verify current pricing and terms with each host before committing.