Ikhtisar benchmark
Definisi dan pemimpin terkini untuk semua 25 benchmark pada 673 model.
GPQA Diamond
๐ง ReasoningGraduate-level science questions in physics, chemistry and biology, designed to be Google-proof.
Humanityโs Last Exam
๐ง ReasoningThousands of expert-written questions across dozens of fields, the hardest closed-ended academic benchmark in wide use.
ARC-AGI-2
๐ง ReasoningSecond-generation abstraction and reasoning tasks measuring fluid intelligence on novel visual patterns.
AIME 2025
๐ MathAmerican Invitational Mathematics Examination problems solved without tool assistance.
FrontierMath
๐ MathResearch-grade mathematics problems crafted by professional mathematicians.
MATH
๐ MathCompetition mathematics dataset spanning algebra to number theory.
SWE-Bench Verified
๐ป CodingReal GitHub issues resolved end-to-end; the industry standard for agentic coding.
SWE-Bench Pro
๐ป CodingHarder, contamination-resistant successor to SWE-Bench with commercial-repo issues.
LiveCodeBench
๐ป CodingContinuously refreshed competitive programming problems immune to data contamination.
SciCode
๐ป CodingScientific coding problems requiring domain knowledge and multi-function solutions.
Terminal-Bench Hard
๐ป CodingMulti-step command-line tasks in realistic terminal environments.
HumanEval
๐ป CodingFunction-level Python completion, the classic code-generation benchmark.
OSWorld
๐ค AgenticReal computer-use tasks across operating systems: browsers, office apps and file management.
Toolathlon
๐ค AgenticMulti-tool orchestration tasks across real-world APIs.
MCP Atlas
๐ค AgenticTool calling over the Model Context Protocol across complex server graphs.
ฯยฒ-Bench Retail
๐ค AgenticDual-control conversational agent tasks in a simulated retail environment.
Apex Agents
๐ค AgenticLong-horizon professional agent tasks with multi-app workflows.
SimpleQA
๐ KnowledgeShort fact-seeking questions measuring hallucination rates.
MMLU-Pro
๐ KnowledgeHarder 14-subject version of MMLU with ten-way multiple choice.
Multilingual MMLU
๐ KnowledgeMMLU translated across 14 languages.
MMMU
๐๏ธ MultimodalMassive Multi-discipline Multimodal Understanding with college-level image reasoning.
MMMU-Pro
๐๏ธ MultimodalHarder multimodal reasoning successor of MMMU.
CharXiv Reasoning
๐๏ธ MultimodalReasoning over scientific charts and figures from arXiv papers.
ScreenSpot-Pro
๐๏ธ MultimodalGUI grounding on high-resolution professional application screens.
MRCR v2
๐ Long ContextMulti-round coreference resolution over needle-in-haystack long contexts.