Why Methodology Matters
With new language models launching every week, a clear evaluation framework is essential. At LLMPodium, we believe rankings should be transparent, reproducible, and grounded in publicly available data.
The Composite Score
Our composite score combines results from multiple benchmarks into a single number between 0 and 100. This gives you a quick way to compare models while still allowing drill-down into individual metrics.
Benchmark Categories
We organize benchmarks into four weighted categories:
Reasoning (35%)
This category tests a model’s ability to think through complex problems. It includes:
- MMLU — 57-subject knowledge test
- GPQA — Graduate-level science questions
- AIME — Competition mathematics
- MATH — Broad mathematical reasoning
Coding (30%)
Code generation and software engineering capabilities:
- HumanEval — Python function synthesis
- SWE-Bench — Real GitHub issue resolution
- LiveCodeBench — Contamination-free competitive programming
Chat (20%)
Instruction following and human preference:
- Arena ELO — Crowd-sourced pairwise preferences
- IFEval — Strict instruction compliance
Multimodal (15%)
Visual and cross-modal reasoning:
- MMMU — College-level multimodal understanding
Normalization
Raw scores are normalized to a 0–100 scale. Percentage-based benchmarks are used directly. ELO scores are scaled relative to the observed range (typically 1100–1400).
Beyond Quality
We also track operational metrics that matter for production:
- Speed — Output tokens per second
- Latency — Time to first token
- Pricing — Cost per million tokens
These don’t affect the composite score but are crucial for real-world deployment decisions.
Transparency
Every data point links to its source. We use publicly available benchmarks, official model cards, and provider pricing pages. We never accept payment for favorable rankings.
What’s Next
We’re working on adding more benchmarks, including agentic task evaluations, long-context tests, and domain-specific assessments. Stay tuned for updates.