Explorateur de Benchmarks

Définitions et leaders actuels pour les 25 benchmarks suivis sur 700 modèles.

GPQA Diamond

reasoning Reasoning

Graduate-level science questions in physics, chemistry and biology, designed to be Google-proof.

🥇GPT-6 Astra96%🥈Claude Mythos Preview94.6%🥉GPT-5.6 Sol94.6%

Humanity’s Last Exam

reasoning Reasoning

Thousands of expert-written questions across dozens of fields, the hardest closed-ended academic benchmark in wide use.

🥇Claude Mythos Preview64.7%🥈Muse Spark58.4%🥉Claude Opus 4.857.9%

ARC-AGI-2

reasoning Reasoning

Second-generation abstraction and reasoning tasks measuring fluid intelligence on novel visual patterns.

🥇GPT-5.585%🥈Gemini 3.1 Pro77.1%🥉GPT-5.473.3%

AIME 2025

math Math

American Invitational Mathematics Examination problems solved without tool assistance.

🥇Grok 4 Heavy100%🥈Gemini 3 Pro100%🥉GPT-5.2100%

FrontierMath

math Math

Research-grade mathematics problems crafted by professional mathematicians.

🥇GPT-6 Astra97.6%🥈GPT-5.6 Sol89%🥉GPT-5.6 Terra84.9%

MATH

math Math

Competition mathematics dataset spanning algebra to number theory.

🥇MiMo V2.5 Pro86.2%🥈GPT-584.7%

SWE-Bench Verified

coding Coding

Real GitHub issues resolved end-to-end; the industry standard for agentic coding.

🥇Claude Fable 595%🥈Claude Mythos Preview93.9%🥉GPT-6 Astra90.2%

SWE-Bench Pro

coding Coding

Harder, contamination-resistant successor to SWE-Bench with commercial-repo issues.

🥇Claude Mythos Preview77.8%🥈Claude Opus 4.869.2%🥉Qwen3.8 Max67.7%

LiveCodeBench

coding Coding

Continuously refreshed competitive programming problems immune to data contamination.

🥇DeepSeek V4 Pro93.5%🥈Gemini 3 Pro91.7%🥉DeepSeek V4 Flash91.6%

SciCode

coding Coding

Scientific coding problems requiring domain knowledge and multi-function solutions.

🥇Claude Fable 560.2%🥈Gemini 3.1 Pro59%🥉Kimi K358.7%

Terminal-Bench Hard

coding Coding

Multi-step command-line tasks in realistic terminal environments.

🥇GPT-6 Astra71.5%🥈GPT-5.6 Sol65.9%🥉Claude Fable 562.9%

HumanEval

coding Coding

Function-level Python completion, the classic code-generation benchmark.

🥇GPT-593.4%

OSWorld

agentic Agentic

Real computer-use tasks across operating systems: browsers, office apps and file management.

🥇Claude Opus 4.672.7%🥈GPT-6 Astra72.6%🥉Claude Sonnet 4.672.5%

Toolathlon

agentic Agentic

Multi-tool orchestration tasks across real-world APIs.

🥇Kimi K373.2%🥈Claude Opus 4.859.9%🥉GPT-5.6 Sol58%

MCP Atlas

agentic Agentic

Tool calling over the Model Context Protocol across complex server graphs.

🥇Kimi K384.2%🥈Claude Opus 4.882.2%🥉Hunyuan Hy379.1%

τ²-Bench Retail

agentic Agentic

Dual-control conversational agent tasks in a simulated retail environment.

🥇GLM-5.299.1%🥈Claude Fable 598.5%🥉Qwen3.6 Plus97.7%

Apex Agents

agentic Agentic

Long-horizon professional agent tasks with multi-app workflows.

🥇Kimi K337.6%🥈Gemini 3.1 Pro33.5%🥉Kimi K2.627.9%

SimpleQA

knowledge Knowledge

Short fact-seeking questions measuring hallucination rates.

🥇Gemini 3 Pro72.1%🥈Gemini 3 Flash68.7%🥉DeepSeek V4 Pro57.9%

MMLU-Pro

knowledge Knowledge

Harder 14-subject version of MMLU with ten-way multiple choice.

🥇Gemini 3 Pro89.8%🥈Qwen3.7 Max89.6%🥉Qwen3.7 Plus88.5%

Multilingual MMLU

knowledge Knowledge

MMLU translated across 14 languages.

🥇Claude Mythos Preview92.7%🥈Gemini 3.1 Pro92.6%🥉Gemini 3 Pro91.8%

MMMU

multimodal Multimodal

Massive Multi-discipline Multimodal Understanding with college-level image reasoning.

🥇Qwen3.6 Plus86%🥈GPT-5.185.4%🥉GPT-584.2%

MMMU-Pro

multimodal Multimodal

Harder multimodal reasoning successor of MMMU.

🥇GPT-6 Astra86.4%🥈GPT-5.583.2%🥉GPT-5.6 Sol83%

CharXiv Reasoning

multimodal Multimodal

Reasoning over scientific charts and figures from arXiv papers.

🥇Claude Mythos Preview93.2%🥈Kimi K391.3%🥉Claude Opus 4.791%

ScreenSpot-Pro

multimodal Multimodal

GUI grounding on high-resolution professional application screens.

🥇Claude Opus 4.887.9%🥈GPT-5.286.3%🥉Qwen3.8 Max84.5%

MRCR v2

long-context Long Context

Multi-round coreference resolution over needle-in-haystack long contexts.

🥇GPT-6 Astra93%🥈GPT-5.6 Sol91.5%🥉GPT-5.6 Terra89.6%
0 / 4