KI-News
Modell-Releases, Benchmark-Updates und Analysen in einem Feed.
Unser jährlicher Daten-Snapshot: Wer führt den Podium Score an, wie groß ist die Open-Weights-Lücke, wo landeten die Preise und welche Benchmarks entschieden das Jahr.
Welche KI-Modelle schreiben 2026 den besten Code? Wir vergleichen Claude Fable 5, Claude Mythos Preview, GPT-5.6 Sol und Kimi K3 bei SWE-Bench Verified, LiveCodeBench und Terminal-Bench.
Kimi K3 erreicht 83.3 im Podium Score — näher an proprietären Frontier-Modellen als je zuvor. Wir quantifizieren die Open-Weights-Lücke 2026 mit aggregierten Ranking-Daten.
Eine Spanne von 178×: die günstigsten und teuersten Frontier-LLMs 2026, Preis-pro-Intelligenz-Analyse und wo der Sweet Spot liegt.
Anthropic's next-generation model family, top-ranked on BenchLM with highest composite score.
Gemini 3.5 Flash-Lite führt mit 389 Tokens/Sekunde. Das komplette Speed-Ranking der Frontier-Modelle und warum Geschwindigkeit der unterschätzteste Benchmark ist.
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.
Compact Inkling variant for fast, cheap inference.
Research preview of Anthropic’s next-generation model family.
Added Alibaba Qwen 3 235B MoE model with hybrid thinking mode to all leaderboards.
Updated LiveCodeBench scores for all models with latest contamination-free results.
Alibaba’s newest hosted flagship, a fast riser on human preference arenas.
Anthropic frontier model with adaptive reasoning effort levels, leading agentic coding benchmarks.
Added Meta Llama 4 Maverick 400B MoE model with 1M context window support.
Lowest-cost Gemini 3.5 tier for extremely high-throughput serving.
Newest Flash-class Gemini, near-pro intelligence at 200+ tokens/s.
Added full benchmark suite for Claude Opus 4 including SWE-Bench and AIME scores.
Moonshot’s frontier MoE with 1M context, top-tier agentic benchmark results.
Thinking Machines’ first frontier model.
Verständnis des Konzepts von LLM-Arenen — Live-Crowdsourcing-Bewertungen, die zeigen, welche Modelle bevorzugt werden.
LLMPodium now available in 10 languages: EN, ZH, JA, KO, TH, RU, DE, ES, IT, FR.
Meta’s proprietary frontier model line, top-10 on blind human preference.
Fast, cheap GPT-5.6 variant built for high-volume production traffic.
Mid-tier GPT-5.6 model balancing intelligence with very high throughput.
Flagship of the GPT-5.6 series for complex reasoning, coding and multi-step agentic workflows.
xAI flagship with real-time knowledge and strong agentic results.
Added Google Gemini 2.5 Pro and 2.5 Flash with thinking capabilities.
Tencent’s latest Hunyuan generation model.
Wie Sie Datenkontamination vermeiden und verlässliche Evaluationen aufbauen.
Updated pricing data for all models from official API documentation.
Balanced Claude tier with adaptive reasoning, strong speed-to-intelligence ratio.
Updated SWE-Bench Verified scores for frontier models.
Added DeepSeek R1 reasoning model with chain-of-thought capabilities.
Wie LLMPodium Ergebnisse mehrerer öffentlicher Ranglisten normalisiert.
New Arena page for side-by-side model comparison launched.
Z.ai flagship with strong agent tool use and 150+ tokens/s.
Code-specialized Kimi variant for repository-level engineering.
Anthropic flagship with adaptive reasoning and Opus-class fallback, tuned for the hardest open-ended and agentic tasks.
NVIDIA’s 550B open MoE built for enterprise reasoning pipelines.
Nex AGI’s efficiency-focused frontier model.