🧮 Best AI for Math Problems
Compare AI models for mathematical problem solving — from basic algebra to competition-level problems and proofs.Rankings combine multiple benchmark scores weighted by relevance to this use case. Updated August 2026.
Why This Matters
Math Problems is one of the most common AI workloads — and the best model depends heavily on the specific task. A model that excels at creative writing may struggle with structured data extraction, and vice versa.
We built this use-case ranking by combining multiple benchmark categories with weights tuned to match real-world usage patterns. For math problems, the score emphasizes the benchmarks that matter most: task accuracy, output quality, and consistency. Speed and cost are factored in but secondary to quality.
All scores are from public independent benchmarks. For a broader view across all tasks, see the overall LLM leaderboard or the expert picks page.
Quick Answer
The best AI models for math problems are:1. Claude Mythos Preview (Anthropic, score: 99.5), 2. Claude Opus 5 (Anthropic, score: 82.3), 3. Claude Fable 5 (Anthropic, score: 81.0).
| # | Model | Provider | Score | Speed | Price (output) |
|---|---|---|---|---|---|
| 1 | Claude Mythos Preview | Anthropic | 99.5 | 80 t/s | $15/M |
| 2 | Claude Opus 5 | Anthropic | 82.3 | 60 t/s | $25/M |
| 3 | Claude Fable 5 | Anthropic | 81.0 | 71 t/s | $50/M |
| 4 | GPT-5.6 Sol | OpenAI | 77.5 | 72 t/s | $30/M |
| 5 | Kimi K3 | Moonshot AI | 77.1 | 37 t/s | $15/M |
| 6 | MiMo V2.5 Pro | Xiaomi | 75.9 | 65 t/s | $0.87/M |
| 7 | Gemini 3.1 Pro | 75.0 | 136 t/s | $12/M | |
| 8 | Claude Opus 4.6 | Anthropic | 74.2 | 50 t/s | $15/M |
| 9 | Qwen3.7 Max | Alibaba | 73.4 | 203 t/s | $7.5/M |
| 10 | Claude Mythos 5 | Anthropic | 71.0 | 50 t/s | $50/M |
| 11 | Grok 4 Heavy | xAI | 68.7 | 50 t/s | $15/M |
| 12 | Claude Opus 4.8 | Anthropic | 68.3 | 50 t/s | $15/M |
| 13 | GLM-5.2 | Z.ai | 68.1 | 164 t/s | $4.29/M |
| 14 | Gemini 3.5 Flash | 66.5 | 267 t/s | $9/M | |
| 15 | Gemini 3.6 Flash | 65.7 | 233 t/s | $7.5/M |