LlamavsPhi

Comparing Meta and Microsoft. Average Podium Score across top models: Llama (67.4) vs Phi (58.4). Flagship showdown: Muse Spark 1.1 (69.9) vs Phi 3 Medium 4k Instruct (60.0).

Llama (Meta)

The global standard in open-weights foundation models from Meta AI.

Top Model: Muse Spark 1.1
Avg Score: 67.4
Licensing: Open Weights Available

Phi (Microsoft)

Highly capable Small Language Models (SLMs) from Microsoft Research.

Avg Score: 58.4
Licensing: Open Weights Available

Top Models Showdown

RankModelLabScoreCodingReasoningSpeedPrice / 1M
#1Muse Spark 1.1Meta69.964.462.7208 t/s$4.25/M
#2Llama 3.1 Nemotron Ultra 253b V1Meta67.050 t/s$0/M
#3Llama 3.1 405b Instruct Bf16Meta67.050 t/s$0/M
#4Llama 3.1 405b Instruct Fp8Meta67.050 t/s$0/M
#5Llama 3.3 Nemotron 49b Super V1Meta66.050 t/s$0/M
#6Phi 3 Medium 4k InstructMicrosoft60.050 t/s$3/M
#7Wizardlm 70bMicrosoft59.050 t/s$3/M
#8Phi 3 Small 8k InstructMicrosoft59.050 t/s$3/M
#9Wizardlm 13bMicrosoft57.050 t/s$3/M
#10Phi 3 Mini 4k Instruct June 2024Microsoft57.050 t/s$3/M
0 / 4 Models Selected