벤치마크 탐색기

25개 모델에서 추적하는 모든 벤치마크의 정의와 상세 정보.

chat

Arena ELO
Chatbot Arena Elo Rating
Crowd-sourced pairwise preference ranking from live human evaluations.
Max: 1600 ELO · chat · LMSYS, 2024
IFEval
Instruction Following Evaluation
Strict instruction-following compliance across formatting and constraint tasks.
Max: 100 % · chat · Zhou et al., 2023

reasoning

MMLU
Massive Multitask Language Understanding
57-subject test of world knowledge and problem solving across STEM, humanities, social sciences.
Max: 100 % · reasoning · Hendrycks et al., 2021
GPQA
Graduate-Level Google-Proof QA
Expert-level questions across physics, chemistry, biology that resist search-engine retrieval.
Max: 100 % · reasoning · Rein et al., 2023
AIME
American Invitational Mathematics Examination
Competition-level mathematics problems requiring multi-step reasoning.
Max: 100 % · reasoning · MAA, 2024
MATH
Mathematics Aptitude Test of Heuristics
High school and competition math problems across algebra, geometry, number theory.
Max: 100 % · reasoning · Hendrycks et al., 2021

code

HumanEval
HumanEval Code Generation
Python function synthesis from docstrings — measures pass@1 accuracy.
Max: 100 % · code · Chen et al., 2021
SWE-Bench
Software Engineering Benchmark
Real-world GitHub issue resolution across popular open-source repositories.
Max: 100 % · code · Jimenez et al., 2024
LiveCodeBench
Live Code Benchmark
Contamination-free code generation from recent competitive programming problems.
Max: 100 % · code · Jain et al., 2024

image

MMMU
Massive Multi-discipline Multimodal Understanding
College-level multimodal reasoning across art, business, science, health, humanities, tech.
Max: 100 % · image · Yue et al., 2024