Обзор бенчмарков
Определения и детали всех бенчмарков, которые мы отслеживаем для 25 моделей.
chat
Arena ELO
Chatbot Arena Elo Rating
Crowd-sourced pairwise preference ranking from live human evaluations.
IFEval
Instruction Following Evaluation
Strict instruction-following compliance across formatting and constraint tasks.
reasoning
MMLU
Massive Multitask Language Understanding
57-subject test of world knowledge and problem solving across STEM, humanities, social sciences.
GPQA
Graduate-Level Google-Proof QA
Expert-level questions across physics, chemistry, biology that resist search-engine retrieval.
AIME
American Invitational Mathematics Examination
Competition-level mathematics problems requiring multi-step reasoning.
MATH
Mathematics Aptitude Test of Heuristics
High school and competition math problems across algebra, geometry, number theory.
code
HumanEval
HumanEval Code Generation
Python function synthesis from docstrings — measures pass@1 accuracy.
SWE-Bench
Software Engineering Benchmark
Real-world GitHub issue resolution across popular open-source repositories.
LiveCodeBench
Live Code Benchmark
Contamination-free code generation from recent competitive programming problems.
image
MMMU
Massive Multi-discipline Multimodal Understanding
College-level multimodal reasoning across art, business, science, health, humanities, tech.