GrokvsPhi
Comparing xAI and Microsoft. Average Podium Score across top models: Grok (73.6) vs Phi (58.4). Flagship showdown: Grok 4.20 Beta1 (74.0) vs Phi 3 Medium 4k Instruct (60.0).
Grok (xAI)
Real-time knowledge and frontier reasoning models from xAI.
Phi (Microsoft)
Highly capable Small Language Models (SLMs) from Microsoft Research.
Top Models Showdown
| Rank | Model | Lab | Score | Coding | Reasoning | Speed | Price / 1M |
|---|---|---|---|---|---|---|---|
| #1 | Grok 4.20 Beta1 | xAI | 74.0 | — | — | 50 t/s | $15/M |
| #2 | Grok 4.20 Beta 0309 Reasoning | xAI | 74.0 | — | — | 50 t/s | $15/M |
| #3 | Grok 4.20 Multi Agent Beta 0309 | xAI | 74.0 | — | — | 50 t/s | $15/M |
| #4 | Grok 4.1 Thinking | xAI | 73.0 | — | — | 50 t/s | $15/M |
| #5 | Grok 4.1 | xAI | 73.0 | — | — | 50 t/s | $15/M |
| #6 | Phi 3 Medium 4k Instruct | Microsoft | 60.0 | — | — | 50 t/s | $3/M |
| #7 | Wizardlm 70b | Microsoft | 59.0 | — | — | 50 t/s | $3/M |
| #8 | Phi 3 Small 8k Instruct | Microsoft | 59.0 | — | — | 50 t/s | $3/M |
| #9 | Wizardlm 13b | Microsoft | 57.0 | — | — | 50 t/s | $3/M |
| #10 | Phi 3 Mini 4k Instruct June 2024 | Microsoft | 57.0 | — | — | 50 t/s | $3/M |
Looking for more ecosystem comparisons? Grok vs Claude · Grok vs ChatGPT / GPT · Grok vs Kimi · Grok vs DeepSeek · Grok vs Gemini · Grok vs Qwen