coding Coding · llm-stats.com
HumanEval
Function-level Python completion, the classic code-generation benchmark.
1
Modelli testati
GPT-5
Modello di punta
93.4%
Punteggio migliore
OpenAI
Provider
| # | Modello | HumanEval |
|---|---|---|
| 1 | GPT-5 OpenAI | 93.4% |
Dati benchmark aggregati da llm-stats.com. Vedi Metodologia.