Skip to main content
LLMPodium
排行榜
🏆 Overall LeaderboardComposite index across all flagship LLMs💻 Coding & SWE-BenchReal-world programming & software agent benchmark🧠 Deep ReasoningGPQA Diamond, AIME & complex logic🤖 Autonomous AgentsOSWorld, Toolathlon & multi-turn execution🔓 Open WeightsDeepSeek, Qwen, Llama & Mistral
竞技场
⚔️ Human Arena BattlesBlind pairwise Elo rankings (700+ models)⚖️ Side-by-Side CompareHead-to-head metric comparison of up to 4 models💰 Cost-per-Task CalculatorEstimate API economics across real workflows🎯 Model RecommenderFind the optimal model for your budget and speed
模型列表
📦 Model DirectoryDetailed profiles, context windows & pricing🏢 AI Labs & ProvidersOpenAI, Anthropic, Google, DeepSeek, Meta📊 Benchmark MatrixEvaluation methodologies and leader tables
nav.resources
📐 Podium MethodologyMathematical aggregation formula explained📰 Changelog & NewsDaily model updates and new evaluations✍️ Research & ArticlesIn-depth AI benchmarking whitepapers❓ Frequently Asked QuestionsCommon questions on scores and ranking📖 LLM GlossaryDefinitions of TTFT, TPS, Elo, MoE & tokens
🇺🇸 EnglishEN🇨🇳 中文ZH🇮🇳 हिन्दीHI🇪🇸 EspañolES🇫🇷 FrançaisFR🇸🇦 العربيةAR🇧🇩 বাংলাBN🇧🇷 PortuguêsPT🇷🇺 РусскийRU🇵🇰 اردوUR🇮🇩 Bahasa IndonesiaID🇩🇪 DeutschDE🇯🇵 日本語JA🇰🇷 한국어KO🇹🇭 ไทยTH🇮🇹 ItalianoIT
排行榜
🏆 排行榜
Overall Composite LeaderboardCoding & Software AgentsDeep Reasoning & MathAutonomous Agent SwarmsOpen Weights & Self-Hosted
⚔️ 竞技场 & Compare
Human Battle Arena (700+ Models)Side-by-Side Model CompareCost-per-Task CalculatorInteractive Model Finder
📦 Directories & Evals
Model Specifications DirectoryAI Labs & Cloud ProvidersBenchmark Methodologies
📖 Research & Info
Scoring MethodologyResearch Blog & ArticlesChangelog & New ModelsFrequently Asked QuestionsTechnical LLM GlossaryAbout LLMPodium
🌐 Language / Язык / 语言
🇺🇸English🇨🇳中文🇮🇳हिन्दी🇪🇸Español🇫🇷Français🇸🇦العربية🇧🇩বাংলা🇧🇷Português🇷🇺Русский🇵🇰اردو🇮🇩Bahasa Indonesia🇩🇪Deutsch🇯🇵日本語🇰🇷한국어🇹🇭ไทย🇮🇹Italiano
Explore Full Leaderboard →
  1. 首页
  2. /最佳模型

🏆 最佳模型

排名基于独立基准评测。分数由多个标准化基准测试计算并归一化到0–100。

💻

最佳编程AI模型

按编程能力排名的语言模型。

🧠

最佳推理AI模型

按推理能力排名的语言模型。

📚

最佳知识AI模型

按知识准确性排名的语言模型。

👁️

最佳视觉AI模型

按视觉理解能力排名的多模态模型。

💰

最佳性价比AI模型

按每美元智能排名的模型。

⚡

最快AI模型

按生成速度排名的模型。

🏆

最智能AI模型

按综合智能排名的模型。

🔢

最佳数学AI模型

按数学推理排名的模型。

📄

最佳长上下文AI模型

按长上下文能力排名的模型。

🤖

最佳AI Agent模型

按代理能力排名的模型。

🛡️

最安全AI模型

按安全性和对齐排名的模型。

📋

最佳指令遵循AI模型

按指令遵循能力排名的模型。

LLMPodium

The definitive open LLM benchmark aggregator and leaderboard. Continuous evaluations across 719+ models and 25+ benchmarks.

Leaderboards synced daily
排行榜
  • 排行榜
  • 💻 编程
  • 🧠 推理
  • 📐 数学
  • 🤖 智能体
  • 🔓 开源权重
  • ⚡ 速度
  • 💰 性价比
分析工具
  • 竞技场 (719+ models)
  • 模型对比 Side-by-Side
  • 基准测试 Matrix
  • 选模型
  • 模型列表 Directory
  • AI 提供商
  • Cost per Task Calculator
  • #1 Ranked Model Profile
技术资源
  • 新闻 Feed
  • 博客文章 & Research
  • Best AI Models 2026
  • Enterprise AI Use Cases
  • 评分方法
  • 术语表
  • 常见问题
  • 关于我们
© 2026 LLMPodium. 保留所有权利。·所有评测数据均定期从公开权威数据源同步。
llms.txtRSS FeedOpen API