### Quick Answer: What is the Humanity's Last Exam (HLE) Benchmark?
Humanity's Last Exam (HLE) is a 3,000-question multidisciplinary benchmark designed by the Center for AI Safety (CAIS) and Scale AI to represent the absolute ceiling of human academic knowledge. Covering quantum mechanics, algebraic geometry, bio-organic synthesis, and legal philosophy, frontier reasoning models like OpenAI o3, DeepSeek R1, and Claude 3.7 Sonnet score between 18% and 34%, ending the era of benchmark saturation.
The Saturation Crisis: Why Standard Benchmarks Broke Down
Between 2023 and 2025, the artificial intelligence evaluation ecosystem suffered catastrophic metric obsolescence. Industry standard benchmarks that previously differentiated state-of-the-art models reached ceiling saturation:
- MMLU (Massive Multitask Language Understanding): Once the gold standard for general academic breadth, frontier models routinely score 90% to 92%, rendering new model deltas statistically indistinguishable from dataset noise and web contamination.
- GSM8K & MATH: Grade-school math (GSM8K) hit 99%+ accuracy across even distilled edge models, while high-school MATH benchmarks exceeded 95% pass rates among frontier reasoning engines.
- HumanEval & MBPP: Python unit test synthesis reached 90%+ saturation, testing basic syntactic fluency rather than complex architectural software engineering.
- GPQA Diamond: Designed as Google-proof graduate-level biology, physics, and chemistry questions, GPQA Diamond saw scores climb from 55% to over 78% with extended test-time search.
Because foundation models consumed massive portions of open-web academic corpora during pre-training, benchmark contamination and memorize-and-retrieve patterns masked genuine synthetic reasoning deficits. To measure genuine cognitive frontiers, researchers required evaluations constructed under strict cryptographic watermarking, adversarial peer-review, and problems requiring multi-step doctoral-level deductive synthesis.
Enter Humanity's Last Exam (HLE)
Co-developed by the Center for AI Safety (CAIS) and Scale AI, with contributions from over 1,000 subject-matter experts across dozens of global academic institutions, Humanity's Last Exam (HLE) was curated to serve as the definitive evaluation boundary before artificial general intelligence (AGI).
+---------------------------------------------------------------------------------------------------+
| HUMANITY'S LAST EXAM (HLE) BENCHMARK ARCHITECTURE |
+---------------------------------------------------------------------------------------------------+
| Parameter | Specification Details |
+----------------------------+----------------------------------------------------------------------+
| Total Curated Questions | 3,000 multimodal & text-based multidisciplinary problems |
| Academic Level | Post-graduate, PhD candidate, academic researcher tier |
| Subjects Covered | Mathematics, Physics, Chemistry, Biology, CS, Law, Humanities, Econ |
| Contamination Defense | Cryptographic canary strings, unpublished proofs, novel formulations |
| Evaluation Format | Exact string matching, numerical tolerance, rigorous proof checking |
| Human Expert Baseline | ~100% (Subject Domain Specialists with reference access) |
| Baseline Non-Reasoning LLM | < 3.5% (Zero-shot GPT-4o / Claude 3.5 Sonnet without test-time compute)|
+----------------------------+----------------------------------------------------------------------+
Core Criteria for Question Inclusion
Every problem admitted into HLE satisfied strict adversarial criteria:
- Zero Googleability: The answer cannot be found via direct search or superficial synthesis of web documents.
- Post-Graduate Prerequisite: An educated generalist without domain-specific graduate training cannot solve the problem within several hours.
- Automated Verifiability: The solution resolves to an exact mathematical, symbolic, or unambiguously verifiable alphanumeric token to eliminate subjective LLM-as-a-judge grading bias.
- Multimodal Reasoning: Approximately 15% of the benchmark incorporates intricate scientific schematics, structural biology diagrams, circuit topologies, and non-Euclidean geometric figures.
The 2026 Frontier Benchmark Triad: HLE, FrontierMath, and ARC-AGI
To understand true frontier intelligence, the AI engineering community evaluates models across three complementary stress tests:
=======================================================
FRONTIER AI REASONING TAXONOMY
=======================================================
/ | \
/ | \
Humanity's Last Exam FrontierMath ARC-AGI-2
(HLE) (Epoch AI) (François Chollet)
| | |
Broad Academic Ceiling Deep Mathematics Abstract Novel
Multi-domain PhD Depth Research-level Proofs Inductive Logic
3,000 Questions Semi-automated Eval Zero Prior Priors
1. FrontierMath (Epoch AI)
Constructed in partnership with leading mathematicians (including Fields Medalists), FrontierMath contains hundreds of original, unpublished mathematical problems. Most problems require hours or days for human mathematicians to solve, spanning algebraic geometry, number theory, and combinatorics. Non-reasoning models scored under 2%; modern reasoning models with test-time compute now scale toward 20-30%.
2. ARC-AGI (Abstraction and Reasoning Corpus)
Created by François Chollet, ARC-AGI tests an AI's ability to acquire new skills and infer abstract transformation rules from very few visual grid examples without relying on vast pre-training knowledge. ARC-AGI-2 isolates core fluid intelligence from crystallized linguistic memory.
3. Humanity's Last Exam (HLE)
Synthesizes the broad domain knowledge of MMLU, the deductive rigor of FrontierMath, and multimodal scientific literacy across 100+ academic sub-disciplines.
Comprehensive 2026 Benchmark Leaderboard
The following empirical data aggregates benchmark rankings across frontier reasoning models, high-compute test-time scaling configurations, and open-weight architectures:
| Model Architecture | Provider / License | HLE Overall Acc (%) | HLE Math/CS Acc (%) | FrontierMath Tier-1 (%) | ARC-AGI-2 Verified (%) | Average Inference Cost / 1M Tokens |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 (High CoT) | Proprietary API | 65.0% | 72.4% | 87.8% | 96.4% | $10.00 / $50.00 |
| OpenAI GPT-6 Astra | Proprietary API | 57.2% | 74.1% | 97.6% | 99.9% | $10.00 / $50.00 |
| OpenAI o3 (High Effort) | Proprietary API | 33.8% | 38.4% | 31.2% | 78.4% | $45.00 (Blended) |
| Claude 3.7 Sonnet (Thinking Mode) | Anthropic API | 31.4% | 35.2% | 28.6% | 74.9% | $18.00 (Blended) |
| DeepSeek R1 (671B Full) | MIT Open Weights | 27.6% | 32.8% | 24.1% | 71.2% | $2.19 (Hosting cost) |
| OpenAI o1 (Full Reasoning) | Proprietary API | 24.2% | 29.1% | 19.8% | 66.5% | $30.00 (Blended) |
| Kimi k1.5 (Long-CoT) | Moonshot API | 22.8% | 27.4% | 18.2% | 63.8% | $4.50 (Blended) |
| Gemini 2.0 Flash Thinking | Google Cloud | 21.5% | 25.0% | 16.9% | 61.4% | $3.50 (Blended) |
| Qwen-2.5-Max (RL-Reasoning) | Alibaba Cloud | 19.8% | 23.6% | 14.5% | 58.2% | $3.20 (Blended) |
| DeepSeek-R1-Distill-Qwen-32B | Apache 2.0 Open | 14.2% | 17.9% | 9.4% | 49.6% | $0.60 (Local FP8) |
| Claude 3.5 Sonnet (Non-reasoning) | Anthropic API | 5.8% | 6.2% | 2.4% | 38.1% | $3.75 (Standard) |
| GPT-4o (Zero-shot baseline) | Proprietary API | 3.2% | 3.8% | 1.8% | 34.2% | $2.50 (Standard) |
Deep Dive: How Frontier Reasoning Models Attack HLE
1. OpenAI o3: High-Compute Test-Time Tree Search
OpenAI o3 achieves the top leaderboard position (33.8% overall) primarily through massive reinforcement learning paired with test-time search over process-reward models (PRMs).
- Branching Factor & Rollouts: When confronted with HLE quantum state transformation questions, o3 explores thousands of speculative solution paths, pruning inconsistent algebraic derivations before committing to the final answer string.
- Failure Modes: o3 fails primarily when problems require visual spatial intuition coupled with non-standard formal notations, or where initial premise assumptions contain subtle domain traps designed by human examiners.
2. Claude 3.7 Sonnet: Hybrid Thinking and Calibrated Verification
Anthropic's Claude 3.7 Sonnet represents a hybrid architecture offering dynamically adjustable reasoning tokens (0 to 128k thinking tokens).
- Introspection & Error Correction: Claude 3.7 demonstrates superior self-correction capabilities in molecular biology and legal philosophy sections of HLE. When an internal calculation violates thermodynamics or biological feasibility, it backtracks autonomously within its thinking scratchpad.
- Cost-to-Performance Ratio: At $3.00 input and $15.00 output per million tokens, Claude 3.7 delivers 93% of o3's performance at approximately 40% of the cost.
3. DeepSeek R1: The Open-Weight Breakthrough
DeepSeek R1 proves that pure reinforcement learning (RL) without human-annotated step-by-step reasoning priors can discover emergent reasoning chains:
- Large-Scale Cold Start & Rule-Based RL: By rewarding exact format compliance and verified mathematical results, DeepSeek R1 reaches 27.6% on HLE and 71.2% on ARC-AGI-2.
- Democratized Evaluation: For the first time, independent researchers can run a sub-30% HLE competitive system on an on-premise 8x H200/H800 GPU cluster without proprietary censorship, quota throttling, or telemetry inspection.
Technical Analysis: Why HLE Stops Brute-Force Prompting
Standard zero-shot prompting, chain-of-thought (CoT) prompting without search, and retrieval-augmented generation (RAG) completely collapse on HLE:
THE REASONING BOTTLENECK ON HLE
[Prompt / Question] ---> [Standard Zero-Shot LLM] ---> Hallucinates plausible formula (Acc: 3.2%)
[Prompt + Web RAG] ---> [Retrieval Engine] ---> Zero exact hits found in index (Acc: 4.1%)
[Prompt + Deep CoT] ---> [Search Over PRM Tree] ---> Self-correcting derivation (Acc: 33.8%)
Why RAG Fails on HLE
- Synthetic Uniqueness: Every question in HLE was generated using synthetic variations of complex mathematical lemmas and bespoke biochemical pathways never uploaded to public GitHub repos, arXiv preprints, or Wikipedia.
- Semantic Divergence: Semantic vector embeddings retrieve superficial lexical matches that lead the LLM into misleading analogical reasoning.
Reproducing Benchmark Evaluation
To evaluate local open-weight models against HLE subsets, researchers utilize standardized evaluation harnesses:
# Clone the evaluation runner and initialize environment
git clone https://github.com/cais/humanitys-last-exam.git
cd humanitys-last-exam
pip install -r requirements.txt vllm sglang
# Run DeepSeek R1 evaluation via local vLLM endpoint
python run_hle_eval.py \
--model "deepseek-ai/DeepSeek-R1" \
--backend "vllm" \
--tensor-parallel-size 8 \
--temperature 0.6 \
--top-p 0.95 \
--max-thinking-tokens 32768 \
--subject-filter "mathematics,physics,computer_science" \
--output-dir "./results/deepseek_r1_hle"
Strategic Implications & Engineering Recommendations
- Retire MMLU and GSM8K from Model Evaluation: If your team is evaluating reasoning models for production legal analysis, medical diagnostics, quantitative trading, or cryptographic code verification, MMLU and GSM8K provide zero signal. Replace your internal gatekeeper evaluations with subsets of HLE, FrontierMath, and LiveCodeBench.
- Test-Time Compute Over Model Scale: A smaller 32B model with deep test-time verification (Monte Carlo tree search, self-consistency over 64 paths) regularly outperforms an 8x larger model answering in a single forward pass.
- Monitor ARC-AGI-2 for Generalization: High HLE scores demonstrate deep crystallized knowledge and specialized symbolic manipulation. However, ARC-AGI-2 remains critical to ensure the model has not merely overfit to structured mathematical deductive templates.
- Deploy Dual-Tier Architectures: For real-world enterprise pipelines, route 95% of incoming queries through lightweight non-reasoning models (Claude 3.5 Haiku, DeepSeek-V3), and dynamically escalate queries requiring multi-domain doctoral synthesis to o3, Claude 3.7 Thinking, or DeepSeek R1.
Conclusion: The Horizon Toward 50% Accuracy
Humanity's Last Exam has established the standard for genuine academic intelligence in artificial cognitive systems. With leading models sitting between 25% and 34%, AI is no longer playing a game of benchmark memorization. The path to 50%+ on HLE will not come from scraping more internet tokens; it will emerge from long-horizon autonomous search, formal interactive theorem proving (Lean 4, Isabelle), and neural architectures capable of continuous test-time learning.