Benchmark

AIME

American Invitational Mathematics Examination

Competition-level mathematics problems requiring multi-step reasoning.

25
Models Tested
Gemini 2.5 Pro
Top Model
92%
Top Score
100 %
Max Possible
#ModelAIME ScoreProviderComposite
192%Google97.5
2
o3
OpenAI
91.6%OpenAI97.1
3
Claude Opus 4
Anthropic
87.4%Anthropic97.8
4
DeepSeek R1
DeepSeek
79.8%DeepSeek92.3
5
o4-mini
OpenAI
79.4%OpenAI91.5
655.3%xAI88.8
755.2%Google86.7
845.3%Meta87.5
940.5%Anthropic90.3
10
DeepSeek V3
DeepSeek
39.2%DeepSeek86.1
1135.6%Alibaba84.5
1220.1%Google80.5
1316%Anthropic88.4
14
GPT-4o
OpenAI
9.3%OpenAI89.2
158.2%Meta79.8
167.8%Meta82.1
176.8%Alibaba76
186.3%Anthropic79.2
19
Mistral Large
Mistral AI
5.6%Mistral AI76.5
205.2%xAI74.8
214.1%Meta74.3
22
Phi-4
Microsoft
3.8%Microsoft64.5
233.2%OpenAI75.8
24
Mixtral 8x22B
Mistral AI
2.1%Mistral AI65.2
251.5%Cohere60.8
Source: MAA, 2024