Benchmark

MMLU

Massive Multitask Language Understanding

57-subject test of world knowledge and problem solving across STEM, humanities, social sciences.

25
Models Tested
Claude Opus 4
Top Model
92.1%
Top Score
100 %
Max Possible
#ModelMMLU ScoreProviderComposite
1
Claude Opus 4
Anthropic
92.1%Anthropic97.8
292%Google97.5
3
o3
OpenAI
91.2%OpenAI97.1
4
DeepSeek R1
DeepSeek
90.8%DeepSeek92.3
590.4%Anthropic90.3
689.5%Meta87.5
789.2%xAI88.8
8
GPT-4o
OpenAI
88.7%OpenAI89.2
988.7%Anthropic88.4
10
DeepSeek V3
DeepSeek
88.5%DeepSeek86.1
11
o4-mini
OpenAI
88.3%OpenAI91.5
1287.8%Alibaba84.5
1387.3%Meta82.1
1486.5%Google86.7
1586%Meta79.8
1685.2%Alibaba76
1785.1%Google80.5
1884.1%Anthropic79.2
19
Mistral Large
Mistral AI
84%Mistral AI76.5
2083.7%Meta74.3
2183.2%xAI74.8
2282%OpenAI75.8
23
Phi-4
Microsoft
78.5%Microsoft64.5
24
Mixtral 8x22B
Mistral AI
77.8%Mistral AI65.2
2575.6%Cohere60.8
Source: Hendrycks et al., 2021