Which model should I use?

Pick the job, set what you're willing to pay, and see what the measured data says. Capability comes from human preference votes on that specific kind of work; price is live API pricing. Nothing here is a blended "intelligence score" we invented.

Required by every agent framework

44 models measured · Arena text / coding

Rows on the frontier are Pareto-optimal: nothing we track is both stronger and cheaper. "Strength vs leader" rescales the board so 100% = parity with the best model on it.

#ModelProviderScoreStrength$/1MContextToolsFrontier
1 Claude Opus 4.6
anthropic/claude-opus-4.6
Anthropic 1534 100% $10.001M tools frontier
2 Claude Opus 5
anthropic/claude-opus-5
Anthropic 1532 99% $10.001M tools frontier
3 Claude Fable 5
anthropic/claude-fable-5
Anthropic 1519 96% $20.001M tools
4 GLM 5.3 Flash
z-ai/glm-5.3-flash
Z.ai (Zhipu) 1516 95% $0.241.31072M tools frontier
5 Gemini 3.8 Flash
google/gemini-3.8-flash
Google 1516 95% $1.501.048576M tools
6 Claude Opus 4.7
anthropic/claude-opus-4.7
Anthropic 1516 95% $10.001M tools
7 Kimi K3
moonshotai/kimi-k3
Moonshot AI 1511 93% $6.001.048576M tools
8 Claude Fable 5.1
anthropic/claude-fable-5.1
Anthropic 1508 93% $20.001M tools
9 Gemini 3.7 Flash
google/gemini-3.7-flash
Google 1504 91% $1.501.048576M tools
10 Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Anthropic 1504 91% $6.001M tools
11 GLM 5.3
z-ai/glm-5.3
Z.ai (Zhipu) 1503 91% $2.151.31072M tools
12 Qwen3.8 Max (0902)
qwen/qwen3.8-max-0902
Qwen (Alibaba) 1498 90% $3.001M tools
13 Qwen3.7 Max
qwen/qwen3.7-max
Qwen (Alibaba) 1498 90% $2.211M tools
14 Gemini 3.5 Flash
google/gemini-3.5-flash
Google 1493 88% $3.381.048576M tools
15 Gemini 3.6 Flash
google/gemini-3.6-flash
Google 1492 88% $1.501.048576M tools
16 GPT-5.6 Sol
openai/gpt-5.6-sol
OpenAI 1490 88% $4.001.05M tools
17 GLM 5.1
z-ai/glm-5.1
Z.ai (Zhipu) 1487 86% $1.49205K tools
18 GPT-5.6 Terra
openai/gpt-5.6-terra
OpenAI 1487 86% $4.501.05M tools
19 Kimi K2.6
moonshotai/kimi-k2.6
Moonshot AI 1487 86% $1.71262K tools
20 Claude Sonnet 5
anthropic/claude-sonnet-5
Anthropic 1485 86% $4.001M tools
21 Kimi K2.5
moonshotai/kimi-k2.5
Moonshot AI 1485 86% $0.90262K tools
22 Claude Opus 4.8
anthropic/claude-opus-4.8
Anthropic 1485 86% $10.001M tools
23 Grok 4.5
x-ai/grok-4.5
xAI 1483 85% $3.00500K tools
24 GLM 5.2
z-ai/glm-5.2
Z.ai (Zhipu) 1479 84% $1.491.048576M tools
25 Qwen3.8 27B
qwen/qwen3.8-27b
Qwen (Alibaba) 1477 84% $1.061M tools
26 MiniMax M3
minimax/minimax-m3
MiniMax 1473 83% $0.531.048576M tools
27 Qwen3.7 Plus
qwen/qwen3.7-plus
Qwen (Alibaba) 1472 82% $0.561M tools
28 DeepSeek V4 Pro 0813
deepseek/deepseek-v4-pro-0813
DeepSeek 1471 82% $1.571.048576M tools
29 DeepSeek V4 Pro 0423
deepseek/deepseek-v4-pro
DeepSeek 1471 82% $1.201.048576M tools
30 GLM 5V Turbo
z-ai/glm-5v-turbo
Z.ai (Zhipu) 1468 81% $1.90203K tools
31 GPT-5.6 Luna
openai/gpt-5.6-luna
OpenAI 1463 80% $0.451.05M tools
32 Mistral Medium 3.5
mistralai/mistral-medium-3-5
Mistral 1461 79% $3.00262K tools
33 GLM 5
z-ai/glm-5
Z.ai (Zhipu) 1461 79% $0.93205K tools
34 DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731
DeepSeek 1456 78% $0.101.31072M tools frontier
35 DeepSeek V4 Flash 0423
deepseek/deepseek-v4-flash
DeepSeek 1456 78% $0.111.048576M tools
36 Gemini 3.5 Flash Lite
google/gemini-3.5-flash-lite
Google 1454 77% $0.851.048576M tools
37 MiniMax M2.7
minimax/minimax-m2.7
MiniMax 1454 77% $0.53205K tools
38 DeepSeek V3.2
deepseek/deepseek-v3.2
DeepSeek 1449 76% $0.30164K tools
39 MiniMax M2.1
minimax/minimax-m2.1
MiniMax 1421 69% $0.53205K tools
40 Grok 4.3
x-ai/grok-4.3
xAI 1415 67% $1.561M tools
41 Kimi K2 Thinking
moonshotai/kimi-k2-thinking
Moonshot AI 1399 63% $1.07262K tools
42 GLM 4.7 Flash
z-ai/glm-4.7-flash
Z.ai (Zhipu) 1385 59% $0.14203K tools
43 MiniMax M2.5
minimax/minimax-m2.5
MiniMax 1379 58% $0.53205K tools
44 MiniMax M2
minimax/minimax-m2
MiniMax 1371 56% $0.45205K tools

Capability: Arena leaderboard dataset (CC-BY-4.0), published 2026-09-02. Pricing: OpenRouter, retrieved 2026-09-10. Blended price = (input×3 + output×1) ÷ 4.

For agents: this whole decision table is one fetch — GET /api/v1/route/{task}.json, or /api/v1/route/all.json for every task at once. No key, open CORS. See API & Bots.

How to read this honestly

  • Capability here is human preference, not correctness. Arena scores say which answer people preferred in head-to-head votes on that category. A model can be preferred and still be wrong.
  • Elo boards and agent boards use different scales. On the Elo boards a dead tie with the leader is a 0.50 win probability; on the Bradley-Terry agent boards parity is 1.0. We rescale both to "% of leader strength" for display and publish the raw numbers in the API.
  • Models with no score on a board are left out of that task rather than given an estimated one.
  • Price is OpenRouter's listed price at retrieval time; going direct to a provider can be cheaper or dearer.