常见问题解答

关于LLMPodium排名、计算方法与数据来源的所有解答。

Every model gets a Podium Score (0–100): 35% arena Elo, 30% benchmark average, 20% Artificial Analysis Intelligence Index and 15% LLM Stats composite — normalized and blended across five independent leaderboards.

We re-sync all five sources weekly. Major model releases trigger an immediate update cycle. Every page shows the last refresh date.

We aggregate five public leaderboards: Arena.ai (arena Elo), Artificial Analysis (intelligence index and performance), LLM Stats, Vellum and LLMBase. Pricing comes from provider documentation.

The Podium Score is a weighted average of four normalized signals: arena Elo (35%), benchmark average (30%), intelligence index (20%) and LLM Stats score (15%). Missing signals are re-weighted across the remaining sources. See our Methodology page for details.

Not all models are evaluated on every benchmark — some benchmarks require specific capabilities (e.g., vision for MMMU) or haven't been run against newer models yet. We show '—' for unavailable data.

Yes! Visit our Compare page to select up to 4 models and see a side-by-side breakdown of all metrics including benchmarks, pricing, speed, and latency.

Arena Elo is a crowd-sourced preference rating from head-to-head human battles: two anonymous models answer the same prompt, people vote, and votes update an Elo-style rating with confidence intervals. It is the single largest signal in our ranking.

We show input and output pricing per million tokens, sourced directly from each provider's official API pricing page. The Value category blends both prices — lower is better.

Absolutely. We track both proprietary and open-weight models from DeepSeek, Alibaba (Qwen), Moonshot AI (Kimi), Z.ai (GLM), MiniMax, Baidu and others. Open-weight models get their own leaderboard at /leaderboard/open.

TTFT measures how long it takes from sending a request to receiving the first token of the response. Lower is better. It's a key metric for interactive applications where responsiveness matters.

Our leaderboard data is sourced from public benchmarks. Check individual benchmark licenses for reuse terms. A machine-readable summary lives at /llms.txt.

When a provider updates a model, we either update the existing entry or create a new one depending on the significance of the change. Retired models keep their last known scores and get a 'legacy' badge.
0 / 4