llm leaderboard

Best LLM for agents

Agentic tool use covers whether a model drives tools to a finished outcome, stays steerable, and recovers when a command fails.

Rumeqo runs these models inside your team rooms. See what each one costs.

92 of 92 ranked models
Ranked models
rankmodelvendorcompositepricein / outbenchmarks%scorescorescorescorescore
1
Claude Fable 5 (high)anthropic/claude-fable-5:high
Anthropic81.7$10.00 / $50.005 of 6 benchmarks12.09.210.81.213.7
2
Claude Opus 5 (high)anthropic/claude-opus-5:high
Anthropic77.6$5.00 / $25.005 of 6 benchmarks12.010.415.21.113.6
3
GPT-5.6 Sol (xhigh)openai/gpt-5.6-sol:xhigh
OpenAI76.4$5.00 / $30.005 of 6 benchmarks10.78.99.81.210.5
4
Claude Opus 5 (max)anthropic/claude-opus-5:max
Anthropic76.1$5.00 / $25.005 of 6 benchmarks11.96.718.11.114.1
5
Claude Opus 4.6anthropic/claude-opus-4.6
Anthropic74.5$5.00 / $25.006 of 6 benchmarks75.6%6.78.25.01.211.1
6
GPT-5.5 (xhigh)openai/gpt-5.5:xhigh
OpenAI74.4$5.00 / $30.005 of 6 benchmarks8.78.94.91.214.4
7
Kimi K3 (max)moonshotai/kimi-k3:max
MoonshotAI73.6$3.00 / $15.005 of 6 benchmarks10.47.515.31.27.9
8
GPT-5.5 (high)openai/gpt-5.5:high
OpenAI72.8$5.00 / $30.005 of 6 benchmarks7.68.73.91.212.9
9
Claude Opus 4.7 (thinking)anthropic/claude-opus-4.7:thinking
Anthropic69.6$5.00 / $25.005 of 6 benchmarks8.27.96.71.112.7
10
Claude Opus 4.7anthropic/claude-opus-4.7
Anthropic69.1$5.00 / $25.005 of 6 benchmarks7.610.35.51.19.9
11
GPT-5.5openai/gpt-5.5
OpenAI68.7$5.00 / $30.005 of 6 benchmarks6.27.03.71.211.5
12
Grok 4.5x-ai/grok-4.5
SpaceXAI68.7$2.00 / $6.005 of 6 benchmarks5.76.45.91.210.8
13
Claude Opus 4.5 (high)anthropic/claude-opus-4.5:high
Anthropic66.7$5.00 / $25.001 of 6 benchmarks76.8%
14
GLM 5.2 (max)z-ai/glm-5.2:max
Z.ai66.5$0.63 / $1.985 of 6 benchmarks6.76.18.41.25.8
15
Gemini 3 Flash Preview (high)google/gemini-3-flash-preview:high
Google65.6$0.50 / $3.001 of 6 benchmarks75.8%
16
MiniMax M2.5 (high)minimax/minimax-m2.5:high
MiniMax65.6$0.22 / $0.901 of 6 benchmarks75.8%
17
GPT-5.4 (high)openai/gpt-5.4:high
OpenAI65.2$2.50 / $15.005 of 6 benchmarks5.06.24.71.29.2
18
Claude Opus 4.5 (medium)anthropic/claude-opus-4.5:medium
Anthropic63.7$5.00 / $25.001 of 6 benchmarks74.4%
19
Gemini 3 Pro Previewgoogle/gemini-3-pro-preview
Google63.01 of 6 benchmarks74.2%
20
Claude Opus 4.8 (thinking)anthropic/claude-opus-4.8:thinking
Anthropic61.8$5.00 / $25.005 of 6 benchmarks9.58.19.1-0.99.1
21
Claude Sonnet 5 (high)anthropic/claude-sonnet-5:high
Anthropic61.8$2.00 / $10.005 of 6 benchmarks7.45.74.01.010.8
22
GPT-5.2-Codexopenai/gpt-5.2-codex
OpenAI61.5$1.75 / $14.001 of 6 benchmarks72.8%
23
GPT-5.2 (high)openai/gpt-5.2:high
OpenAI61.5$1.75 / $14.001 of 6 benchmarks72.8%
24
GLM 5 (high)z-ai/glm-5:high
Z.ai61.5$0.95 / $2.551 of 6 benchmarks72.8%
25
GPT-5.6 Luna (xhigh)openai/gpt-5.6-luna:xhigh
OpenAI61.5$0.10 / $0.605 of 6 benchmarks4.31.7-0.91.211.3
26
DeepSeek V4 Flash 0423 (high)deepseek/deepseek-v4-flash:high
DeepSeek61.0$0.14 / $0.285 of 6 benchmarks3.92.89.11.24.1
27
Claude Sonnet 4.5 (high)anthropic/claude-sonnet-4.5:high
Anthropic60.0$3.00 / $15.001 of 6 benchmarks71.4%
28
Kimi K2.5 (high)moonshotai/kimi-k2.5:high
MoonshotAI59.3$0.57 / $2.851 of 6 benchmarks70.8%
29
GPT-5.6 Terra (xhigh)openai/gpt-5.6-terra:xhigh
OpenAI59.0$1.00 / $6.005 of 6 benchmarks3.45.1-1.31.29.3
30
Claude Sonnet 4.5anthropic/claude-sonnet-4.5
Anthropic58.6$3.00 / $15.001 of 6 benchmarks70.6%
31
DeepSeek V3.2 (high)deepseek/deepseek-v3.2:high
DeepSeek57.8$0.269 / $0.401 of 6 benchmarks70.0%
32
Claude Opus 4.8anthropic/claude-opus-4.8
Anthropic57.8$5.00 / $25.005 of 6 benchmarks2.58.28.7-28.610.5
33
Gemini 3 Pro Preview (high)google/gemini-3-pro-preview:high
Google57.11 of 6 benchmarks69.6%
34
GPT-5.2openai/gpt-5.2
OpenAI56.3$1.75 / $14.001 of 6 benchmarks69.0%
35
Claude Sonnet 4.6anthropic/claude-sonnet-4.6
Anthropic55.9$3.00 / $15.005 of 6 benchmarks3.02.2-0.21.110.9
36
Claude Opus 4anthropic/claude-opus-4
Anthropic55.6$15.00 / $75.001 of 6 benchmarks67.6%
37
Muse Spark 1.1meta/muse-spark-1.1
Meta55.3$1.25 / $4.255 of 6 benchmarks1.1-3.26.71.15.7
38
Claude Haiku 4.5 (high)anthropic/claude-haiku-4.5:high
Anthropic54.8$1.00 / $5.001 of 6 benchmarks66.6%
39
GPT-5.1-Codex (medium)openai/gpt-5.1-codex:medium
OpenAI53.7$1.25 / $10.001 of 6 benchmarks66.0%
40
GPT-5.1 (medium)openai/gpt-5.1:medium
OpenAI53.7$1.25 / $10.001 of 6 benchmarks66.0%
41
Kimi K2.7 Codemoonshotai/kimi-k2.7-code
MoonshotAI53.1$0.67 / $3.405 of 6 benchmarks1.0-2.04.61.2-1.6
42
GPT-5 (medium)openai/gpt-5:medium
OpenAI52.6$1.25 / $10.001 of 6 benchmarks65.0%
43
Claude Sonnet 4anthropic/claude-sonnet-4
Anthropic51.9$3.00 / $15.001 of 6 benchmarks64.9%
44
Kimi K2 Thinkingmoonshotai/kimi-k2-thinking
MoonshotAI51.1$0.60 / $2.501 of 6 benchmarks63.4%
45
MiniMax M2minimax/minimax-m2
MiniMax50.4$0.255 / $1.021 of 6 benchmarks61.0%
46
DeepSeek V3.2 (thinking)deepseek/deepseek-v3.2:thinking
DeepSeek49.7$0.269 / $0.401 of 6 benchmarks60.0%
47
Kimi K2.6moonshotai/kimi-k2.6
MoonshotAI49.1$0.95 / $4.005 of 6 benchmarks-0.6-1.00.71.2-5.8
48
GPT-5 Mini (medium)openai/gpt-5-mini:medium
OpenAI48.9$0.25 / $2.001 of 6 benchmarks59.8%
49
o3openai/o3
OpenAI48.2$2.00 / $8.001 of 6 benchmarks58.4%
50
Devstral Small 2512mistralai/devstral-small-2512
Mistral AI47.41 of 6 benchmarks56.4%
51
GPT-5 Miniopenai/gpt-5-mini
OpenAI46.7$0.25 / $2.001 of 6 benchmarks56.2%
52
Qwen3.7 Maxqwen/qwen3.7-max
Qwen46.6$1.475 / $4.4255 of 6 benchmarks-0.0-0.2-0.10.54.7
53
Qwen3 Coder 480B A35b Instructqwen/qwen3-coder-480b-a35b-instruct
Qwen45.61 of 6 benchmarks55.4%
54
GLM 4.6z-ai/glm-4.6
Z.ai45.6$0.50 / $2.001 of 6 benchmarks55.4%
55
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview
Google45.5$2.00 / $12.005 of 6 benchmarks-0.63.11.40.9-11.6
56
GLM 4.5z-ai/glm-4.5
Z.ai44.5$0.60 / $2.201 of 6 benchmarks54.2%
57
Devstral 2mistralai/devstral-2
Mistral AI43.71 of 6 benchmarks53.8%
58
GLM 5.1z-ai/glm-5.1
Z.ai43.5$1.40 / $4.405 of 6 benchmarks0.51.51.7-0.5-1.0
59
Gemini 2.5 Progoogle/gemini-2.5-pro
Google43.0$1.25 / $10.001 of 6 benchmarks53.6%
60
DeepSeek V4 Prodeepseek/deepseek-v4-pro
DeepSeek42.7$1.168 / $2.3365 of 6 benchmarks-0.10.6-2.20.34.0
61
Claude 3.7 Sonnetanthropic/claude-3-7-sonnet
Anthropic42.31 of 6 benchmarks52.8%
62
Gemini 3.5 Flash (high)google/gemini-3.5-flash:high
Google41.6$1.50 / $9.005 of 6 benchmarks-0.4-0.70.50.3-1.9
63
o4 Miniopenai/o4-mini
OpenAI41.5$1.10 / $4.401 of 6 benchmarks45.0%
64
Kimi K2 Instructmoonshotai/kimi-k2-instruct
Moonshot AI40.81 of 6 benchmarks43.8%
65
Gemini 3.6 Flashgoogle/gemini-3.6-flash
Google40.1$1.50 / $7.505 of 6 benchmarks-2.8-5.0-0.81.1-3.6
66
GPT-4.1openai/gpt-4.1
OpenAI40.0$2.00 / $8.001 of 6 benchmarks39.6%
67
Qwen3.7 Plusqwen/qwen3.7-plus
Qwen39.3$0.32 / $1.285 of 6 benchmarks-1.8-4.9-1.00.25.9
68
GPT-5 Nano (medium)openai/gpt-5-nano:medium
OpenAI39.3$0.05 / $0.401 of 6 benchmarks34.8%
69
Gemini 2.5 Flashgoogle/gemini-2.5-flash
Google38.6$0.30 / $2.501 of 6 benchmarks28.7%
70
MiniMax M3minimax/minimax-m3
MiniMax38.2$0.30 / $1.205 of 6 benchmarks-2.5-4.8-5.80.65.5
71
gpt-oss-120bopenai/gpt-oss-120b
OpenAI37.8$0.03 / $0.171 of 6 benchmarks26.0%
72
GPT-4.1 Miniopenai/gpt-4.1-mini
OpenAI37.1$0.40 / $1.601 of 6 benchmarks23.9%
73
GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20
OpenAI36.3$2.50 / $10.001 of 6 benchmarks21.6%
74
MiMo-V2.5-Proxiaomi/mimo-v2.5-pro
Xiaomi35.7$0.435 / $0.875 of 6 benchmarks-2.2-2.4-3.2-0.11.7
75
Llama 4 Maverick 17B 128e Instructmeta-llama/llama-4-maverick-17b-128e-instruct
Meta35.61 of 6 benchmarks21.0%
76
DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash
DeepSeek35.6$0.14 / $0.285 of 6 benchmarks-2.3-1.2-2.2-1.32.5
77
Gemini 2.0 Flash 001google/gemini-2.0-flash-001
Google34.81 of 6 benchmarks13.5%
78
Llama 4 Scout 17B 16e Instructmeta-llama/llama-4-scout-17b-16e-instruct
Meta34.11 of 6 benchmarks9.1%
79
Gemini 3.5 Flash (medium)google/gemini-3.5-flash:medium
Google33.6$1.50 / $9.005 of 6 benchmarks-3.6-3.9-8.20.4-0.6
80
Qwen2.5 Coder 32B Instructqwen/qwen2.5-coder-32b-instruct
Qwen33.41 of 6 benchmarks9.0%
81
Hy3tencent/hy3
Tencent33.2$0.132 / $0.5285 of 6 benchmarks-1.3-8.5-2.4-1.33.5
82
Inklingthinkingmachines/inkling
Thinking Machines32.3$0.95 / $4.055 of 6 benchmarks-6.7-11.8-12.50.56.4
83
Grok 4.3 (high)x-ai/grok-4.3:high
SpaceXAI30.4$1.25 / $2.505 of 6 benchmarks-8.5-7.0-9.61.0-13.0
84
Grok Build 0.1x-ai/grok-build-0.1
SpaceXAI29.0$1.00 / $2.005 of 6 benchmarks-9.0-8.8-5.60.9-20.7
85
Grok 4.3x-ai/grok-4.3
SpaceXAI27.7$1.25 / $2.505 of 6 benchmarks-14.6-6.0-10.91.1-41.1
86
Gemini 3 Flash Previewgoogle/gemini-3-flash-preview
Google27.5$0.50 / $3.005 of 6 benchmarks-8.6-3.9-7.50.2-21.1
87
MiniMax M2.7minimax/minimax-m2.7
MiniMax25.2$0.30 / $1.205 of 6 benchmarks-11.1-13.5-11.31.0-16.7
88
Mistral Medium 3.5mistralai/mistral-medium-3-5
Mistral25.2$1.50 / $7.505 of 6 benchmarks-6.9-11.2-9.4-3.0-1.2
89
Solar Pro 4upstage/solar-pro4
Upstage23.6$0.03 / $0.125 of 6 benchmarks-12.1-13.6-7.30.3-16.5
90
Gemma 4 31Bgoogle/gemma-4-31b-it
Google23.0$0.10 / $0.345 of 6 benchmarks-18.2-9.30.1-30.9-48.6
91
Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite
Google22.4$0.30 / $2.505 of 6 benchmarks-10.2-9.9-13.9-0.4-13.0
92
Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b
NVIDIA20.2$0.60 / $3.605 of 6 benchmarks-14.6-19.7-15.80.5-23.7
How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model scored on fewer than 3 of the 6 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.

Data sources

Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.

  • LMArenaCC BY 4.0

    Arena ratings by LMArena, from the public leaderboard dataset.

  • SWE-bench

    Resolve rates published by the SWE-bench maintainers.