AI 新闻

模型发布、基准更新与分析——来自我们的数据管道。

2026-08-05
文章
2026年LLM现状:领奖台报告

我们的年度数据快照:谁领跑 Podium Score、开放权重差距有多大、价格落在哪里,以及哪些基准测试决定了这一年。

2026-08-03
文章
2026年编程最佳LLM:SWE-Bench、LiveCodeBench与Terminal-Bench领跑者

2026年哪些AI模型的代码写得最好?我们对比 Claude Fable 5、Claude Mythos Preview、GPT-5.6 Sol 与 Kimi K3 在 SWE-Bench Verified、LiveCodeBench 和 Terminal-Bench 上的表现。

2026-08-02
文章
2026年开放权重 vs 专有LLM:差距有多大?

Kimi K3 的 Podium Score 达到 83.3——比以往任何时候都更接近前沿专有模型。我们用排行榜汇总数据量化2026年的开放权重差距。

2026-08-01
文章
LLM定价对比(2026):每百万输出 token 从 $0.28 到 $50

178倍的价差:2026年最便宜与最贵的前沿LLM、每点智能的价格分析,以及价值甜区在哪里。

2026-08-01
发布
Anthropic releases Claude Mythos 5

Anthropic's next-generation model family, top-ranked on BenchLM with highest composite score.

2026-07-30
文章
2026年最快的LLM:输出速度与延迟对比

Gemini 3.5 Flash-Lite 以 389 token/秒领跑。前沿模型的完整速度排名,以及为什么速度是最被低估的基准。

2026-07-30
发布
DeepSeek releases DeepSeek V4 Flash

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.

2026-07-29
发布
Thinking Machines releases Inkling Small

Compact Inkling variant for fast, cheap inference.

2026-07-29
发布
Anthropic releases Claude Mythos Preview

Research preview of Anthropic’s next-generation model family.

2026-07-28
模型
Added Qwen 3 235B

Added Alibaba Qwen 3 235B MoE model with hybrid thinking mode to all leaderboards.

2026-07-25
基准
LiveCodeBench scores updated

Updated LiveCodeBench scores for all models with latest contamination-free results.

2026-07-24
发布
Alibaba releases Qwen3.8 Max

Alibaba’s newest hosted flagship, a fast riser on human preference arenas.

2026-07-23
发布
Anthropic releases Claude Opus 5

Anthropic frontier model with adaptive reasoning effort levels, leading agentic coding benchmarks.

2026-07-20
模型
Added Llama 4 Maverick

Added Meta Llama 4 Maverick 400B MoE model with 1M context window support.

2026-07-20
发布
Google releases Gemini 3.5 Flash-Lite

Lowest-cost Gemini 3.5 tier for extremely high-throughput serving.

2026-07-20
发布
Google releases Gemini 3.6 Flash

Newest Flash-class Gemini, near-pro intelligence at 200+ tokens/s.

2026-07-15
更新
Claude Opus 4 benchmark data

Added full benchmark suite for Claude Opus 4 including SWE-Bench and AIME scores.

2026-07-15
发布
Moonshot AI releases Kimi K3

Moonshot’s frontier MoE with 1M context, top-tier agentic benchmark results.

2026-07-14
发布
Thinking Machines releases Inkling

Thinking Machines’ first frontier model.

2026-07-10
文章
什么是LLM竞技场?为什么它如此重要?

深入了解LLM竞技场概念——通过实时众包评估揭示人们真正偏好的AI模型。

2026-07-10
功能
Multi-language support

LLMPodium now available in 10 languages: EN, ZH, JA, KO, TH, RU, DE, ES, IT, FR.

2026-07-08
发布
Meta releases Muse Spark 1.1

Meta’s proprietary frontier model line, top-10 on blind human preference.

2026-07-08
发布
OpenAI releases GPT-5.6 Luna

Fast, cheap GPT-5.6 variant built for high-volume production traffic.

2026-07-08
发布
OpenAI releases GPT-5.6 Terra

Mid-tier GPT-5.6 model balancing intelligence with very high throughput.

2026-07-08
发布
OpenAI releases GPT-5.6 Sol

Flagship of the GPT-5.6 series for complex reasoning, coding and multi-step agentic workflows.

2026-07-07
发布
xAI releases Grok 4.5

xAI flagship with real-time knowledge and strong agentic results.

2026-07-05
模型
Added Gemini 2.5 Pro/Flash

Added Google Gemini 2.5 Pro and 2.5 Flash with thinking capabilities.

2026-07-05
发布
Tencent releases Hunyuan Hy3

Tencent’s latest Hunyuan generation model.

2026-07-01
文章
2026年LLM基准测试最佳实践

如何防止数据污染、应对提示词敏感性并构建可靠的AI模型评估。

2026-07-01
更新
Pricing data refresh

Updated pricing data for all models from official API documentation.

2026-06-29
发布
Anthropic releases Claude Sonnet 5

Balanced Claude tier with adaptive reasoning, strong speed-to-intelligence ratio.

2026-06-28
基准
SWE-Bench Verified scores

Updated SWE-Bench Verified scores for frontier models.

2026-06-20
模型
Added DeepSeek R1

Added DeepSeek R1 reasoning model with chain-of-thought capabilities.

2026-06-15
文章
理解LLM评估方法论

深入分析LLMPodium如何跨多个公开排行榜归一化并加权计算综合得分。

2026-06-15
功能
Arena page launched

New Arena page for side-by-side model comparison launched.

2026-06-15
发布
Z.ai releases GLM-5.2

Z.ai flagship with strong agent tool use and 150+ tokens/s.

2026-06-11
发布
Moonshot AI releases Kimi K2.7 Code

Code-specialized Kimi variant for repository-level engineering.

2026-06-08
发布
Anthropic releases Claude Fable 5

Anthropic flagship with adaptive reasoning and Opus-class fallback, tuned for the hardest open-ended and agentic tasks.

2026-06-03
发布
NVIDIA releases Nemotron 3 Ultra 550B

NVIDIA’s 550B open MoE built for enterprise reasoning pipelines.

2026-06-01
发布
Nex AGI releases Nex N2 Pro

Nex AGI’s efficiency-focused frontier model.