Voice & Audio Intelligence
Speech-to-Text & Audio AI Arena
Standardized evaluation across the AA-WER v2 composite index (AgentTalk, VoxPopuli, Earnings22), real-time factor (RTF), and streaming latency.
| # | Model & Provider | AA-WER Error Rate (Lower is Better) | Speed Factor (RTF) | Time to First Partial | Cost / 1,000 Minutes | Streaming Support |
|---|---|---|---|---|---|---|
| 1 | Nova-3 (Speech-to-Text) Deepgram | 4.8% | 110x speed | 0.18s | $4.30 | Streaming & Batch |
| 2 | Universal-3.5-Pro AssemblyAI | 5.1% | 95x speed | 0.22s | $6.50 | Streaming & Batch |
| 3 | Whisper large-v3 turbo OpenAI / Open Source | 5.4% | 85x speed | 0.35s | $6.00 | Batch Async |
| 4 | Gemini 3.5 Transcribe Pro Google Cloud | 5.2% | 75x speed | 0.25s | $7.20 | Streaming & Batch |
| 5 | Scribe v1 (Conversational) ElevenLabs | 5.6% | 60x speed | 0.28s | $8.00 | Streaming & Batch |