💻 Coding · llm-stats.com
SWE-Bench Verified
Real GitHub issues resolved end-to-end; the industry standard for agentic coding.
28
Model diuji
Claude Fable 5
Model teratas
95%
Skor terbaik
Anthropic
Penyedia
| # | Model | SWE-Bench Verified |
|---|---|---|
| 1 | Claude Fable 5 Anthropic | 95% |
| 2 | Claude Mythos Preview Anthropic | 93.9% |
| 3 | Claude Opus 4.8 Anthropic | 88.6% |
| 4 | Claude Opus 4.7 Anthropic | 87.6% |
| 5 | Claude Sonnet 5 Anthropic | 85.2% |
| 6 | Claude Opus 4.5 Anthropic | 80.9% |
| 7 | Claude Opus 4.6 Anthropic | 80.8% |
| 8 | Gemini 3.1 Pro Google | 80.6% |
| 9 | DeepSeek V4 Pro DeepSeek | 80.6% |
| 10 | MiniMax M3 MiniMax | 80.5% |
| 11 | Qwen3.7 Max Alibaba | 80.4% |
| 12 | Kimi K2.6 Moonshot AI | 80.2% |
| 13 | GPT-5.2 OpenAI | 80% |
| 14 | Claude Sonnet 4.6 Anthropic | 79.6% |
| 15 | DeepSeek V4 Flash DeepSeek | 79% |
| 16 | MiMo V2.5 Pro Xiaomi | 78.9% |
| 17 | Qwen3.6 Plus Alibaba | 78.8% |
| 18 | Gemini 3 Flash Google | 78% |
| 19 | Hunyuan Hy3 Tencent | 78% |
| 20 | Qwen3.7 Plus Alibaba | 77.7% |
| 21 | Muse Spark Meta | 77.4% |
| 22 | Seed 2.0 Pro ByteDance | 76.5% |
| 23 | Qwen3.5 397B-A17B Alibaba | 76.4% |
| 24 | GPT-5.1 OpenAI | 76.3% |
| 25 | Gemini 3 Pro Google | 76.2% |
| 26 | GPT-5 OpenAI | 74.9% |
| 27 | Claude Opus 4.1 Anthropic | 74.5% |
| 28 | DeepSeek V3.2 DeepSeek | 73.1% |
Data benchmark digabungkan dari llm-stats.com. Selengkapnya di metodologi.