Benchmarks

Penyedia API LLM Termurah & Tercepat 2026: Groq vs Cerebras vs DeepInfra

Jawaban Cepat: Pada 2026, Cerebras adalah API LLM tercepat mencapai 450-920 tok/s pada Llama 3.3 70B, sementara Groq menghadirkan 280-310 tok/s dengan TTFT sub-150ms. Sebaliknya, DeepInfra adalah penyedia API LLM termurah yang menawarkan DeepSeek V3 seharga $0.14/$0.28 per 1M token. Untuk keandalan tinggi, Fireworks dan OpenRouter menyediakan failover otomatis.

1. Ringkasan Eksekutif: Lanskap Inferensi LLM Tahun 2026

Ekonomi kecerdasan buatan generatif telah berubah drastis pada tahun 2026. Dua paradigma perangkat keras baru telah menggantikan monopoli pusat data GPU konvensional:

  1. Arsitektur ASIC Ultra-Cepat dengan SRAM On-Chip: Penyedia seperti Cerebras (Wafer-Scale Engine WSE-3) dan Groq (Language Processing Unit LPU) mengatasi hambatan memori HBM dengan menyimpan bobot model sepenuhnya di SRAM on-chip, menghasilkan kecepatan 250 hingga 900+ token per detik (tok/s).
  2. Klaster GPU Optimal Berbiaya Rendah (vLLM / TensorRT-LLM): DeepInfra, Together AI, dan Fireworks AI mengoperasikan klaster Nvidia H100 SXM5 dan H200 dengan continuous batching dan prefix caching, menekan harga per juta token ke titik terendah.
  3. Gateway Perutean Cerdas: Platform seperti OpenRouter mendistribusikan beban kerja ke lebih dari 40 pusat data dengan failover otomatis tanpa gangguan.

Tim LLMPodium menguji Groq, Cerebras, DeepInfra, Together AI, Fireworks AI, dan OpenRouter menggunakan Meta Llama 3.3 70B Instruct, DeepSeek V3 (671B MoE), dan Alibaba Qwen 2.5 72B Instruct.


2. Matriks Benchmark Komprehensif: Kecepatan, Latensi, dan Biaya

+-----------------------------------------------------------------------------------------------------------------------------------------+
|                                    MATRIKS BENCHMARK PENYEDIA API LLM 2026 (LLAMA 3.3 70B & DEEPSEEK V3)                                |
+------------------+-------------------+-----------------+--------------------+--------------------+--------------------+-------------+
| Penyedia         | Mesin Perangkat   | TTFT (P50/P95)  | Output TPS (70B)   | Input $/1M Token   | Output $/1M Token  | Uptime SLA  |
+------------------+-------------------+-----------------+--------------------+--------------------+--------------------+-------------+
| Cerebras         | WSE-3 (Wafer ASIC)| 112ms / 185ms   | 485 - 920 tok/s    | $0.60              | $0.60              | 99.9%       |
| Groq             | LPU (Custom TSP)  | 138ms / 215ms   | 280 - 315 tok/s    | $0.59              | $0.79              | 99.9%       |
| Fireworks AI     | FireAttention H100| 185ms / 310ms   | 135 - 165 tok/s    | $0.90              | $0.90              | 99.95%      |
| Together AI      | TurboEngine H100  | 210ms / 340ms   | 115 - 145 tok/s    | $0.88              | $0.88              | 99.9%       |
| DeepInfra        | vLLM / H100 SXM5  | 310ms / 540ms   | 68 - 88 tok/s      | $0.23              | $0.40              | 99.8%       |
| OpenRouter (Auto)| Multi-Provider GW | 245ms / 520ms   | 75 - 320 tok/s     | $0.23 - $0.90      | $0.40 - $0.90      | 99.99%*     |
+------------------+-------------------+-----------------+--------------------+--------------------+--------------------+-------------+

3. Analisis Perangkat Keras: Groq vs Cerebras

+----------------------------------------------------------------------------------------------+
|                                    PERBANDINGAN GROQ VS CEREBRAS                             |
+--------------------------+---------------------------------+---------------------------------+
| Fitur Arsitektur         | Groq LPU                        | Cerebras WSE-3                  |
+--------------------------+---------------------------------+---------------------------------+
| Kecepatan Output (70B)   | 280 - 315 tok/s                 | 485 - 920 tok/s (1.8x - 2.9x)   |
| Latensi TTFT (P50)       | 138 ms                          | 112 ms                          |
| Tarif (Llama 3.3 70B)    | $0.59 in / $0.79 out per 1M     | $0.60 in / $0.60 out per 1M     |
| Window Konteks           | 128k token                      | 8k - 32k token                  |
| Tool Calling & JSON      | Standar produksi (100% schema)  | Berkembang cepat (98.2%)        |
| Ragam Model              | Llama 3.3, Qwen 2.5, DeepSeek   | Llama 3.3 70B, Llama 3.1 8B     |
+--------------------------+---------------------------------+---------------------------------+
  • Groq API Key: Tersedia langsung di Groq Cloud Console dengan dukungan OpenAI SDK dan kuota gratis 30 RPM.
  • DeepInfra: Paling hemat biaya untuk DeepSeek V3 ($0.14/$0.28 per 1 juta token).
  • Fireworks AI & Together AI: Sangat handal dengan SLA 99.95% dan akurasi JSON schema tinggi.

4. Hasil Pengujian pada 3 Model Terkemuka

+-------------------------------------------------------------------------------------------------------------------------+
|                                  PERBANDINGAN KECEPATAN DAN BIAYA PADA 3 MODEL FONDASI                                  |
+-----------------------+----------------------+--------------------+--------------------+-------------------+------------+
| Model                 | Penyedia             | Rata-rata TPS      | TTFT (P50)         | Total Biaya 1M    | Error Rate |
+-----------------------+----------------------+--------------------+--------------------+-------------------+------------+
| Llama 3.3 70B         | Cerebras             | 620 tok/s          | 112 ms             | $1.20 (total)     | 0.04%      |
| Llama 3.3 70B         | Groq                 | 295 tok/s          | 138 ms             | $1.38 (total)     | 0.02%      |
| Llama 3.3 70B         | Fireworks AI         | 148 tok/s          | 185 ms             | $1.80 (total)     | 0.01%      |
| Llama 3.3 70B         | Together AI          | 126 tok/s          | 210 ms             | $1.76 (total)     | 0.03%      |
| Llama 3.3 70B         | DeepInfra            | 78 tok/s           | 310 ms             | $0.63 (total)     | 0.06%      |
+-----------------------+----------------------+--------------------+--------------------+-------------------+------------+
| DeepSeek V3 (671B MoE)| DeepInfra            | 82 tok/s           | 285 ms             | $0.42 (total)     | 0.05%      |
| DeepSeek V3 (671B MoE)| Fireworks AI         | 115 tok/s          | 220 ms             | $1.80 (total)     | 0.02%      |
| DeepSeek V3 (671B MoE)| Together AI          | 98 tok/s           | 240 ms             | $2.00 (total)     | 0.03%      |
| DeepSeek V3 (671B MoE)| OpenRouter (Auto)    | 88 tok/s           | 295 ms             | $0.42 - $1.80     | 0.00%*     |
+-----------------------+----------------------+--------------------+--------------------+-------------------+------------+
| Qwen 2.5 72B Instruct | DeepInfra            | 74 tok/s           | 320 ms             | $0.63 (total)     | 0.04%      |
| Qwen 2.5 72B Instruct | Fireworks AI         | 138 tok/s          | 195 ms             | $1.80 (total)     | 0.02%      |
| Qwen 2.5 72B Instruct | Together AI          | 118 tok/s          | 225 ms             | $1.76 (total)     | 0.03%      |
| Qwen 2.5 72B Instruct | Groq                 | 280 tok/s          | 145 ms             | $1.38 (total)     | 0.02%      |
+-----------------------+----------------------+--------------------+--------------------+-------------------+------------+

5. Panduan Implementasi Produksi

# Uji latensi Groq API Key
export GROQ_API_KEY="gsk_your_groq_api_key_here"

curl -X POST "https://api.groq.com/openai/v1/chat/completions"   -H "Authorization: Bearer $GROQ_API_KEY"   -H "Content-Type: application/json"   -d '{
    "model": "llama-3.3-70b-versatile",
    "messages": [{"role": "user", "content": "Jelaskan speculative decoding dalam 2 kalimat."}],
    "stream": true
  }'   -w "
Koneksi: %{time_connect}s | TTFT: %{time_starttransfer}s | Total: %{time_total}s
"
import os
import time
from openai import OpenAI

class ResilientLLMClient:
    def __init__(self):
        self.fast_client = OpenAI(
            base_url="https://api.groq.com/openai/v1",
            api_key=os.environ.get("GROQ_API_KEY")
        )
        self.cheap_client = OpenAI(
            base_url="https://api.deepinfra.com/v1/openai",
            api_key=os.environ.get("DEEPINFRA_TOKEN")
        )

    def generate(self, prompt: str, prioritize_speed: bool = True) -> str:
        if prioritize_speed:
            try:
                start = time.perf_counter()
                res = self.fast_client.chat.completions.create(
                    model="llama-3.3-70b-versatile",
                    messages=[{"role": "user", "content": prompt}],
                    max_tokens=800
                )
                print(f"[Groq LPU] Waktu: {time.perf_counter() - start:.3f}s")
                return res.choices[0].message.content
            except Exception as e:
                print(f"[Groq Peringatan] Beralih ke DeepInfra: {e}")

        start = time.perf_counter()
        res = self.cheap_client.chat.completions.create(
            model="deepseek-ai/DeepSeek-V3",
            messages=[{"role": "user", "content": prompt}],
            max_tokens=800
        )
        print(f"[DeepInfra] Waktu: {time.perf_counter() - start:.3f}s")
        return res.choices[0].message.content

6. Simulasi Biaya Bulanan: 100M Token

  • Claude 3.5 Sonnet: $675.00/bulan
  • Cerebras (Llama 3.3 70B): $75.00/bulan (hemat 88%)
  • Groq (Llama 3.3 70B): $78.75/bulan (hemat 88%)
  • DeepInfra (DeepSeek V3): $21.00/bulan (hemat 97%)

7. Rekomendasi Strategis

  1. Aplikasi Suara & Chat Real-Time: Cerebras untuk throughput tertinggi; Groq untuk konteks 128k dan tool calling.
  2. Pemrosesan Batch & RAG Skala Besar: DeepInfra dengan DeepSeek V3 untuk efisiensi biaya terbaik.
  3. Keandalan Sistem Enterprise: Fireworks AI atau arsitektur multi-endpoint dengan OpenRouter.
← Semua artikel
0 / 4