بینچ مارک کا جائزہ

673 ماڈلز پر تمام 25 بینچ مارکس کی تعریفیں اور موجودہ سرکردہ۔

GPQA Diamond

🧠 Reasoning

Graduate-level science questions in physics, chemistry and biology, designed to be Google-proof.

🥇Claude Mythos Preview94.6%🥈GPT-5.6 Sol94.6%🥉Gemini 3.1 Pro94.3%

Thousands of expert-written questions across dozens of fields, the hardest closed-ended academic benchmark in wide use.

🥇Claude Mythos Preview64.7%🥈Muse Spark58.4%🥉Claude Opus 4.857.9%

ARC-AGI-2

🧠 Reasoning

Second-generation abstraction and reasoning tasks measuring fluid intelligence on novel visual patterns.

🥇GPT-5.585%🥈Gemini 3.1 Pro77.1%🥉GPT-5.473.3%

AIME 2025

📐 Math

American Invitational Mathematics Examination problems solved without tool assistance.

🥇Grok 4 Heavy100%🥈Gemini 3 Pro100%🥉GPT-5.2100%

FrontierMath

📐 Math

Research-grade mathematics problems crafted by professional mathematicians.

🥇GPT-5.6 Sol89%🥈GPT-5.6 Terra84.9%🥉GPT-5.6 Luna78.6%

MATH

📐 Math

Competition mathematics dataset spanning algebra to number theory.

🥇MiMo V2.5 Pro86.2%🥈GPT-584.7%

Real GitHub issues resolved end-to-end; the industry standard for agentic coding.

🥇Claude Fable 595%🥈Claude Mythos Preview93.9%🥉Claude Opus 4.888.6%

SWE-Bench Pro

💻 Coding

Harder, contamination-resistant successor to SWE-Bench with commercial-repo issues.

🥇Claude Mythos Preview77.8%🥈Claude Opus 4.869.2%🥉Qwen3.8 Max67.7%

LiveCodeBench

💻 Coding

Continuously refreshed competitive programming problems immune to data contamination.

🥇DeepSeek V4 Pro93.5%🥈Gemini 3 Pro91.7%🥉DeepSeek V4 Flash91.6%

SciCode

💻 Coding

Scientific coding problems requiring domain knowledge and multi-function solutions.

🥇Claude Fable 560.2%🥈Gemini 3.1 Pro59%🥉Kimi K358.7%

Multi-step command-line tasks in realistic terminal environments.

🥇GPT-5.6 Sol65.9%🥈Claude Fable 562.9%🥉GPT-5.560.6%

HumanEval

💻 Coding

Function-level Python completion, the classic code-generation benchmark.

🥇GPT-593.4%

OSWorld

🤖 Agentic

Real computer-use tasks across operating systems: browsers, office apps and file management.

🥇Claude Opus 4.672.7%🥈Claude Sonnet 4.672.5%

Toolathlon

🤖 Agentic

Multi-tool orchestration tasks across real-world APIs.

🥇Kimi K373.2%🥈Claude Opus 4.859.9%🥉GPT-5.6 Sol58%

MCP Atlas

🤖 Agentic

Tool calling over the Model Context Protocol across complex server graphs.

🥇Kimi K384.2%🥈Claude Opus 4.882.2%🥉Hunyuan Hy379.1%

τ²-Bench Retail

🤖 Agentic

Dual-control conversational agent tasks in a simulated retail environment.

🥇GLM-5.299.1%🥈Claude Fable 598.5%🥉Qwen3.6 Plus97.7%

Apex Agents

🤖 Agentic

Long-horizon professional agent tasks with multi-app workflows.

🥇Kimi K337.6%🥈Gemini 3.1 Pro33.5%🥉Kimi K2.627.9%

SimpleQA

📚 Knowledge

Short fact-seeking questions measuring hallucination rates.

🥇Gemini 3 Pro72.1%🥈Gemini 3 Flash68.7%🥉DeepSeek V4 Pro57.9%

MMLU-Pro

📚 Knowledge

Harder 14-subject version of MMLU with ten-way multiple choice.

🥇Gemini 3 Pro89.8%🥈Qwen3.7 Max89.6%🥉Qwen3.7 Plus88.5%

Multilingual MMLU

📚 Knowledge

MMLU translated across 14 languages.

🥇Claude Mythos Preview92.7%🥈Gemini 3.1 Pro92.6%🥉Gemini 3 Pro91.8%

MMMU

👁️ Multimodal

Massive Multi-discipline Multimodal Understanding with college-level image reasoning.

🥇Qwen3.6 Plus86%🥈GPT-5.185.4%🥉GPT-584.2%

MMMU-Pro

👁️ Multimodal

Harder multimodal reasoning successor of MMMU.

🥇GPT-5.583.2%🥈GPT-5.6 Sol83%🥉Qwen3.8 Max82.3%

CharXiv Reasoning

👁️ Multimodal

Reasoning over scientific charts and figures from arXiv papers.

🥇Claude Mythos Preview93.2%🥈Kimi K391.3%🥉Claude Opus 4.791%

ScreenSpot-Pro

👁️ Multimodal

GUI grounding on high-resolution professional application screens.

🥇Claude Opus 4.887.9%🥈GPT-5.286.3%🥉Qwen3.8 Max84.5%

MRCR v2

📜 Long Context

Multi-round coreference resolution over needle-in-haystack long contexts.

🥇GPT-5.6 Sol91.5%🥈GPT-5.6 Terra89.6%🥉Claude Opus 4.676%