coding Coding · llm-stats.com

HumanEval

Function-level Python completion, the classic code-generation benchmark.

1
Modelli testati
GPT-5
Modello di punta
93.4%
Punteggio migliore
OpenAI
Provider
#ModelloHumanEvalProviderPodium Score
1
GPT-5
OpenAI
93.4%
OpenAI23.0
Dati benchmark aggregati da llm-stats.com. Vedi Metodologia.
0 / 4