Clasificación
Clasificación: Investigación
Literatura científica y conocimiento experto: HLE, GPQA Diamond y MMLU-Pro.
General700Código399Matemáticas53Razonamiento413Agéntico55Conocimiento412Multimodal382Agentes de código55Razonamiento científico82Multilingüe12Escritura61Chat y asistentes61Investigación85Llamada de herramientas32Contexto largo82Velocidad402Valor401Pesos abiertos170
85 modelos
| # | Modelo | Podium Score | Velocidad |
|---|---|---|---|
| 01 | Seed 2.0 Pro ByteDance | 88.9 | 40 t/s |
| 02 | GPT-5.1 OpenAI | 88.1 | 60 t/s |
| 03 | Claude Opus 4.5 Anthropic | 87.0 | 50 t/s |
| 4 | ERNIE 5.1 Baidu | 87.0 | 60 t/s |
| 5 | Qwen3.5 122B-A10B Alibaba · Abierto | 86.7 | 126 t/s |
| 6 | GPT-5 OpenAI | 85.7 | 60 t/s |
| 7 | Gemma 4 31B Google · Abierto | 85.2 | 35 t/s |
| 8 | DeepSeek V3.2 DeepSeek · Abierto | 85.0 | 60 t/s |
| 9 | Qwen3.8 Flash Next Alibaba · Abierto | 81.0 | 168 t/s |
| 10 | Claude Mythos Preview Anthropic · reasoning · preview | 79.7 | 80 t/s |
| 11 | GLM-5.3 Flash Z.ai · Abierto | 78.5 | 182 t/s |
| 12 | Claude Opus 4.8 Anthropic | 75.8 | 50 t/s |
| 13 | Gemini 3 Pro Google | 75.8 | 120 t/s |
| 14 | DeepSeek V4 Pro DeepSeek · Abierto | 75.3 | 66 t/s |
| 15 | Kimi K3 Moonshot AI · Abierto · reasoning | 74.8 | 37 t/s |
| 16 | GPT-6 Astra OpenAI · reasoning | 74.6 | 82 t/s |
| 17 | Claude Opus 4.7 Anthropic | 74.5 | 80 t/s |
| 18 | Qwen3.7 Max Alibaba | 74.5 | 203 t/s |
| 19 | Muse Spark Meta | 74.0 | 80 t/s |
| 20 | DeepSeek V4 Flash DeepSeek · Abierto | 73.1 | 113 t/s |
| 21 | GLM-5.2 Z.ai · Abierto | 73.0 | 164 t/s |
| 22 | Claude Fable 5 Anthropic · reasoning | 72.9 | 71 t/s |
| 23 | Claude Opus 5 Anthropic | 72.9 | 60 t/s |
| 24 | Gemini 3.1 Pro Google | 72.9 | 136 t/s |
| 25 | GPT-5.5 OpenAI | 72.9 | 60 t/s |
| 26 | Claude Opus 4.6 Anthropic | 72.2 | 50 t/s |
| 27 | Qwen3.7 Plus Alibaba | 71.2 | 54 t/s |
| 28 | GPT-5.6 Sol OpenAI | 70.9 | 72 t/s |
| 29 | Grok 4 Heavy xAI · preview | 69.6 | 50 t/s |
| 30 | Claude Sonnet 4.6 Anthropic | 69.5 | 50 t/s |
| 31 | Qwen3.6 Plus Alibaba | 69.2 | 55 t/s |
| 32 | Qwen3.5 397B-A17B Alibaba · Abierto | 68.3 | 71 t/s |
| 33 | Qwen3.8 Max Alibaba | 68.1 | 55 t/s |
| 34 | Granite 4.2 30B Instruct IBM · Abierto | 68.0 | 78 t/s |
| 35 | Muse Spark 1.1 Meta | 67.5 | 208 t/s |
| 36 | GPT-5.6 Terra OpenAI | 67.4 | 127 t/s |
| 37 | GLM-5.1 Z.ai · Abierto | 67.0 | 74 t/s |
| 38 | Gemini 3 Flash Google | 67.0 | 120 t/s |
| 39 | Claude Sonnet 4.6 (max) Anthropic | 67.0 | 50 t/s |
| 40 | Nemotron 3 Ultra 550B NVIDIA · Abierto | 66.8 | 142 t/s |
| 41 | Grok 4.5 xAI | 66.7 | 59 t/s |
| 42 | Gemini 3.5 Flash Google | 66.6 | 267 t/s |
| 43 | GPT-5.4 OpenAI | 66.3 | 60 t/s |
| 44 | GPT-5.3 Codex OpenAI | 65.7 | 118 t/s |
| 45 | Gemini 3.6 Flash Google | 65.6 | 233 t/s |
| 46 | MiniMax M3 MiniMax | 65.0 | 93 t/s |
| 47 | GPT-5.6 Luna OpenAI | 64.8 | 175 t/s |
| 48 | MiniMax-M2.7 MiniMax | 64.0 | 40 t/s |
| 49 | GPT-5.2 OpenAI | 63.5 | 90 t/s |
| 50 | Kimi K2.6 Moonshot AI · Abierto | 63.5 | 40 t/s |
| 51 | DeepSeek V4.1 Flash DeepSeek · Abierto · reasoning | 62.0 | 214 t/s |
| 52 | Kimi K2.7 Code Moonshot AI · Abierto | 61.2 | 39 t/s |
| 53 | Hunyuan Hy3 Tencent | 61.0 | 69 t/s |
| 54 | MiMo V2.5 Pro Xiaomi · Abierto | 60.2 | 65 t/s |
| 55 | Inkling Thinking Machines | 58.5 | 81 t/s |
| 56 | Granite 4.2 8B Instruct IBM · Abierto | 54.0 | 115 t/s |
| 57 | Mistral Medium 3.5 Mistral | 51.0 | 60 t/s |
| 58 | Agnes 3.0 Flash Agnes · reasoning | 49.0 | 235 t/s |
| 59 | Claude Sonnet 5 Anthropic | 48.9 | 96 t/s |
| 60 | Claude 4.5 Haiku Anthropic | 48.0 | 80 t/s |
| 61 | Nova 2.0 Pro Preview Amazon · preview | 46.0 | 40 t/s |
| 62 | Agnes 2.5 Pro Beta Agnes | 45.0 | 65 t/s |
| 63 | GPT-oss-120B OpenAI · Abierto | 43.0 | 240 t/s |
| 64 | Granite 4.2 3B Instruct IBM · Abierto | 42.0 | 148 t/s |
| 65 | DeepSeek V4 Flash Vision DeepSeek | 36.0 | 120 t/s |
| 66 | G9v3-39A5B AI9Stars · Abierto | 36.0 | 60 t/s |
| 67 | G9v3-3B AI9Stars · Abierto | 36.0 | 60 t/s |
| 68 | GLM-5.3 Zhipu AI | 36.0 | 90 t/s |
| 69 | 36.0 | 59 t/s | |
| 70 | 36.0 | 59 t/s | |
| 71 | 36.0 | 62 t/s | |
| 72 | HyperNova 60B Multiverse · Abierto | 36.0 | 367 t/s |
| 73 | KAT-Coder-Pro V2 Kwai | 36.0 | 103 t/s |
| 74 | LFM2.5-2.6B Liquid AI · Abierto | 36.0 | 60 t/s |
| 75 | LFM2.5-8B-A1B Liquid AI · Abierto | 36.0 | 338 t/s |
| 76 | LFM2.5-VL-1.6B Liquid AI · Abierto | 36.0 | 395 t/s |
| 77 | LongCat 2.0 LongCat · Abierto | 36.0 | 42 t/s |
| 78 | Magistral Small 1.2 Mistral · Abierto | 36.0 | 60 t/s |
| 79 | Qwen3.8 27B (xhigh) Alibaba · Abierto | 36.0 | 53 t/s |
| 80 | Qwen3.8 27B (low) Alibaba · Abierto | 36.0 | 64 t/s |
| 81 | Qwen3.8 27B (medium) Alibaba · Abierto | 36.0 | 62 t/s |
| 82 | Qwen3.8 27B (Non-reasoning) Alibaba · Abierto · reasoning | 36.0 | 62 t/s |
| 83 | Solar Pro 3 Upstage | 34.0 | 40 t/s |
| 84 | GPT-oss-20B OpenAI · Abierto | 32.0 | 90 t/s |
| 85 | 31.0 | 40 t/s |
Podium Score es una mezcla ponderada de cinco clasificaciones. Consulta la Metodología para más detalles.