👁️ Multimodal · llm-stats.com

CharXiv Reasoning

Reasoning over scientific charts and figures from arXiv papers.

12
Modelos testados
Claude Mythos Preview
Modelo principal
93.2%
Melhor pontuação
Anthropic
Provedor
#ModeloCharXiv ReasoningProvedorPodium Score
1
93.2%
Anthropic97.4
2
Kimi K3
Moonshot AI
91.3%
Moonshot AI83.3
3
91%
Anthropic52.0
4
89.9%
Anthropic76.0
5
Kimi K2.6
Moonshot AI
86.7%
Moonshot AI49.0
6
86.4%
Meta47.0
7
85.9%
Alibaba42.2
8
GPT-5.2
OpenAI
82.1%
OpenAI49.2
9
81.5%
Alibaba32.4
10
81.4%
Google51.9
11
80.3%
Google42.4
12
77.4%
Anthropic61.2
Dados de benchmark agregados de llm-stats.com. Saiba mais na metodologia.