Independent
analysis of AI
Understand the frontier AI landscape to choose the optimal foundation model, pricing tier, and multi-cloud inference provider for your use case.
Highlights
Intelligence Quality
Output Speed
Cost per 1M Tokens
Top Ranked Models
Top 10 Models by Overall Podium Score
View All →| # | Model | Podium Score |
|---|---|---|
| 01 | Claude Mythos Preview Anthropic · Proprietary | 97.4 |
| 02 | Claude Fable 5 Anthropic · Proprietary | 93.4 |
| 03 | Claude Fable 5.1 Anthropic · Proprietary | 88.4 |
| 4 | GPT-6 Astra OpenAI · Proprietary | 86.2 |
| 5 | Kimi K3 Moonshot AI · Open Weights | 83.3 |
| 6 | Claude Opus 5 Anthropic · Proprietary | 81.4 |
| 7 | GPT-5.6 Sol OpenAI · Proprietary | 80.8 |
| 8 | Qwen3.8 Max Alibaba · Proprietary | 79.8 |
| 9 | Claude Opus 4.8 Anthropic · Proprietary | 76.0 |
| 10 | Claude Opus 4 6 Thinking Anthropic · Unknown | 75.0 |
Best by Use Case
Coding
Repository-level software engineering: SWE-Bench Verified and Pro, LiveCodeBench, SciCode, Terminal-Bench and HumanEval.
Math
Mathematical problem solving from competition math to frontier research problems: AIME, MATH and FrontierMath.
Reasoning
Deep reasoning and scientific knowledge: GPQA Diamond, Humanity’s Last Exam and ARC-AGI-2.
Agentic
Autonomous tool use and computer operation: OSWorld, Toolathlon, MCP Atlas, τ²-Bench Retail and Apex Agents.
Knowledge
Factual knowledge and hallucination resistance: SimpleQA, MMLU-Pro and Multilingual MMLU.
Multimodal
Vision and document understanding: MMMU, MMMU-Pro, CharXiv Reasoning and ScreenSpot-Pro.