Frequently Asked Questions
Everything you need to know about LLMPodium rankings, methodology, and data.
We compute a composite score from multiple benchmark results including MMLU, GPQA, SWE-Bench, HumanEval, AIME, and Arena ELO. The composite is a weighted average emphasizing real-world task performance across reasoning, coding, and instruction following.
We update benchmark scores and pricing data on a weekly basis. Major model releases trigger an immediate update cycle. Each data point is timestamped so you can see freshness.
Scores come from official benchmark leaderboards (LMSYS Chatbot Arena, Papers With Code, official model cards) and independent evaluations. Pricing data comes directly from provider APIs.
The composite score is a weighted average of all tracked benchmark results for a model. Weights are calibrated to reflect real-world utility — coding and reasoning benchmarks carry more weight than pure knowledge tests. See our Methodology page for exact weights.
Not all models are evaluated on every benchmark. Some benchmarks require specific capabilities (e.g., vision for MMMU) or haven't been run against newer models yet. We show '—' for unavailable data.
Yes! Visit our Compare page to select up to 4 models and see a side-by-side breakdown of all metrics including benchmarks, pricing, speed, and latency.
Arena ELO is a crowd-sourced preference ranking from LMSYS Chatbot Arena, where humans compare model outputs side-by-side in blind tests. It's one of the most reliable measures of perceived model quality.
We show input and output pricing per million tokens, sourced directly from each provider's official API pricing page. For models available through multiple providers, we use the most common or cheapest option.
Absolutely. We track both proprietary and open-weight models. Open-weight models are tagged with an 'Open' badge in the leaderboard. We include models from Meta (Llama), Mistral, DeepSeek, Alibaba (Qwen), and others.
TTFT measures how long it takes from sending a request to receiving the first token of the response. Lower is better. It's a key metric for interactive applications where responsiveness matters.
Our leaderboard data is sourced from public benchmarks. Check individual benchmark licenses for reuse terms. We plan to offer a public API for programmatic access in the future.
When a provider updates a model (e.g., GPT-4o → GPT-4o-2025), we either update the existing entry or create a new one depending on the significance of the change. Version dates are tracked.