Changelog
Track all updates to LLMPodium: new models, benchmark updates, and feature releases.
Added Alibaba Qwen 3 235B MoE model with hybrid thinking mode to all leaderboards.
Updated LiveCodeBench scores for all models with latest contamination-free results.
Added Meta Llama 4 Maverick 400B MoE model with 1M context window support.
Added full benchmark suite for Claude Opus 4 including SWE-Bench and AIME scores.
LLMPodium now available in 10 languages: EN, ZH, JA, KO, TH, RU, DE, ES, IT, FR.
Added Google Gemini 2.5 Pro and 2.5 Flash with thinking capabilities.
Updated pricing data for all models from official API documentation.
Updated SWE-Bench Verified scores for frontier models.
Added DeepSeek R1 reasoning model with chain-of-thought capabilities.
New Arena page for side-by-side model comparison launched.