Benchmark Explorer
Definitions and current leaders for all 25 benchmarks tracked across 700 models.
GPQA Diamond
reasoning ReasoningGraduate-level science questions in physics, chemistry and biology, designed to be Google-proof.
Humanity’s Last Exam
reasoning ReasoningThousands of expert-written questions across dozens of fields, the hardest closed-ended academic benchmark in wide use.
ARC-AGI-2
reasoning ReasoningSecond-generation abstraction and reasoning tasks measuring fluid intelligence on novel visual patterns.
AIME 2025
math MathAmerican Invitational Mathematics Examination problems solved without tool assistance.
FrontierMath
math MathResearch-grade mathematics problems crafted by professional mathematicians.
MATH
math MathCompetition mathematics dataset spanning algebra to number theory.
SWE-Bench Verified
coding CodingReal GitHub issues resolved end-to-end; the industry standard for agentic coding.
SWE-Bench Pro
coding CodingHarder, contamination-resistant successor to SWE-Bench with commercial-repo issues.
LiveCodeBench
coding CodingContinuously refreshed competitive programming problems immune to data contamination.
SciCode
coding CodingScientific coding problems requiring domain knowledge and multi-function solutions.
Terminal-Bench Hard
coding CodingMulti-step command-line tasks in realistic terminal environments.
HumanEval
coding CodingFunction-level Python completion, the classic code-generation benchmark.
OSWorld
agentic AgenticReal computer-use tasks across operating systems: browsers, office apps and file management.
Toolathlon
agentic AgenticMulti-tool orchestration tasks across real-world APIs.
MCP Atlas
agentic AgenticTool calling over the Model Context Protocol across complex server graphs.
τ²-Bench Retail
agentic AgenticDual-control conversational agent tasks in a simulated retail environment.
Apex Agents
agentic AgenticLong-horizon professional agent tasks with multi-app workflows.
SimpleQA
knowledge KnowledgeShort fact-seeking questions measuring hallucination rates.
MMLU-Pro
knowledge KnowledgeHarder 14-subject version of MMLU with ten-way multiple choice.
Multilingual MMLU
knowledge KnowledgeMMLU translated across 14 languages.
MMMU
multimodal MultimodalMassive Multi-discipline Multimodal Understanding with college-level image reasoning.
MMMU-Pro
multimodal MultimodalHarder multimodal reasoning successor of MMMU.
CharXiv Reasoning
multimodal MultimodalReasoning over scientific charts and figures from arXiv papers.
ScreenSpot-Pro
multimodal MultimodalGUI grounding on high-resolution professional application screens.
MRCR v2
long-context Long ContextMulti-round coreference resolution over needle-in-haystack long contexts.