Benchmark

MMMU

Massive Multi-discipline Multimodal Understanding

College-level multimodal reasoning across art, business, science, health, humanities, tech.

25
Models Tested
o3
Top Model
80.4%
Top Score
100 %
Max Possible
#ModelMMMU ScoreProviderComposite
1
o3
OpenAI
80.4%OpenAI97.1
279.5%Google97.5
3
Claude Opus 4
Anthropic
78.2%Anthropic97.8
4
o4-mini
OpenAI
74.8%OpenAI91.5
573.1%Meta87.5
672.8%xAI88.8
772.5%Anthropic90.3
8
DeepSeek R1
DeepSeek
71.5%DeepSeek92.3
970.1%Google86.7
10
GPT-4o
OpenAI
69.1%OpenAI89.2
1168.3%Anthropic88.4
1266.5%Alibaba84.5
1364.5%Meta82.1
1464.3%Google80.5
15
DeepSeek V3
DeepSeek
64.2%DeepSeek86.1
1662.8%Meta79.8
1761.2%Alibaba76
1860.2%Anthropic79.2
19
Mistral Large
Mistral AI
59.8%Mistral AI76.5
2059.4%OpenAI75.8
2158.5%xAI74.8
2258.2%Meta74.3
23
Phi-4
Microsoft
52.3%Microsoft64.5
24
Mixtral 8x22B
Mistral AI
48.5%Mistral AI65.2
2545.1%Cohere60.8
Source: Yue et al., 2024