Changelog
Track all updates to LLMPodium: new models, benchmark updates, and feature releases.
Added Sapiens AI Agnes 3.0 Flash (234 t/s, 1M context) and DeepSeek V4.1 Flash (552B MoE, 16B active, MIT open-weights reasoning).
Added OpenAI GPT-6 Astra frontier reasoning model with 1.1M context, 100% ExploitBench, and Critical-tier autonomous cybersecurity safeguards.
Added Alibaba Qwen 3 235B MoE model with hybrid thinking mode to all leaderboards.
Updated LiveCodeBench scores for all models with latest contamination-free results.
Added Meta Llama 4 Maverick 400B MoE model with 1M context window support.
Added full benchmark suite for Claude Opus 4 including SWE-Bench and AIME scores.
LLMPodium now available in 16 languages: EN, ZH, JA, KO, TH, RU, DE, ES, IT, FR, HI, AR, BN, PT, UR, ID.
Added Google Gemini 2.5 Pro and 2.5 Flash with thinking capabilities.
Updated pricing data for all models from official API documentation.
Updated SWE-Bench Verified scores for frontier models.
Added DeepSeek R1 reasoning model with chain-of-thought capabilities.
New Arena page for side-by-side model comparison launched.