Berita AI
Rilis model, pembaruan benchmark, dan analisis dari pipeline data kami.
Our annual data snapshot: who leads the Podium Score, how wide the open-weights gap is, where prices landed, and which benchmarks decided the year.
Which AI models write the best code in 2026? We compare Claude Fable 5, Claude Mythos Preview, GPT-5.6 Sol and Kimi K3 across SWE-Bench Verified, LiveCodeBench and Terminal-Bench.
Kimi K3 scores 83.3 on the Podium Score — closer to frontier proprietary models than ever. We quantify the 2026 open-weights gap using aggregated leaderboard data.
A 178× spread: the cheapest and most expensive frontier LLMs in 2026, price-per-intelligence analysis, and where the value sweet spot is.
Anthropic's next-generation model family, top-ranked on BenchLM with highest composite score.
Gemini 3.5 Flash-Lite leads at 389 tokens/second. Full speed ranking of frontier models, and why speed is the most underrated benchmark.
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.
Compact Inkling variant for fast, cheap inference.
Research preview of Anthropic’s next-generation model family.
Added Alibaba Qwen 3 235B MoE model with hybrid thinking mode to all leaderboards.
Updated LiveCodeBench scores for all models with latest contamination-free results.
Alibaba’s newest hosted flagship, a fast riser on human preference arenas.
Anthropic frontier model with adaptive reasoning effort levels, leading agentic coding benchmarks.
Added Meta Llama 4 Maverick 400B MoE model with 1M context window support.
Lowest-cost Gemini 3.5 tier for extremely high-throughput serving.
Newest Flash-class Gemini, near-pro intelligence at 200+ tokens/s.
Added full benchmark suite for Claude Opus 4 including SWE-Bench and AIME scores.
Moonshot’s frontier MoE with 1M context, top-tier agentic benchmark results.
Thinking Machines’ first frontier model.
Understanding the concept of LLM arenas — live, crowd-sourced evaluations that reveal which models people actually prefer.
LLMPodium now available in 10 languages: EN, ZH, JA, KO, TH, RU, DE, ES, IT, FR.
Meta’s proprietary frontier model line, top-10 on blind human preference.
Fast, cheap GPT-5.6 variant built for high-volume production traffic.
Mid-tier GPT-5.6 model balancing intelligence with very high throughput.
Flagship of the GPT-5.6 series for complex reasoning, coding and multi-step agentic workflows.
xAI flagship with real-time knowledge and strong agentic results.
Added Google Gemini 2.5 Pro and 2.5 Flash with thinking capabilities.
Tencent’s latest Hunyuan generation model.
How to avoid data contamination, account for prompt sensitivity, and build reliable evaluations.
Updated pricing data for all models from official API documentation.
Balanced Claude tier with adaptive reasoning, strong speed-to-intelligence ratio.
Updated SWE-Bench Verified scores for frontier models.
Added DeepSeek R1 reasoning model with chain-of-thought capabilities.
Deep dive into how LLMPodium normalizes and weights scores across multiple public leaderboards.
New Arena page for side-by-side model comparison launched.
Z.ai flagship with strong agent tool use and 150+ tokens/s.
Code-specialized Kimi variant for repository-level engineering.
Anthropic flagship with adaptive reasoning and Opus-class fallback, tuned for the hardest open-ended and agentic tasks.
NVIDIA’s 550B open MoE built for enterprise reasoning pipelines.
Nex AGI’s efficiency-focused frontier model.