Best AI model for tool-call reliability

How rarely the model hallucinates or misuses tool calls during agent runs. Higher is better. Ranked by Arena agent_tool_hallucination, priced against live API rates. 29 models have a published score on this board.

Capability published 2026-09-08 · pricing retrieved 2026-09-10

Most capable
GPT-6 Astra
OpenAI
$20.00 / 1M blended
100% of leader strength
openai/gpt-6-astra
Best value
DeepSeek V4 Flash 0731
DeepSeek
$0.10 / 1M blended
92% of leader strength · 100% cheaper
deepseek/deepseek-v4-flash-0731
Budget
DeepSeek V4 Flash 0731
DeepSeek
$0.10 / 1M blended
92% of leader strength · 100% cheaper
deepseek/deepseek-v4-flash-0731
Best for agents
DeepSeek V4 Flash 0731
DeepSeek
$0.10 / 1M blended
92% of leader strength · 100% cheaper
deepseek/deepseek-v4-flash-0731

Top 25 models for tool-call reliability

Rows marked frontier are Pareto-optimal — nothing we track is both stronger and cheaper. "Strength" rescales the board so 100% = parity with the leader.

#ModelProviderScoreStrength$/1MContextTools
1 GPT-6 Astra
openai/gpt-6-astra
OpenAI 0.0039 100% $20.001.05M tools
2 Claude Fable 5.1
anthropic/claude-fable-5.1
Anthropic 0.0039 100% $20.001M tools
3 GLM 5.3 Flash
z-ai/glm-5.3-flash
Z.ai (Zhipu) 0.0039 100% $0.241.31072M tools
4 GLM 5.3
z-ai/glm-5.3
Z.ai (Zhipu) 0.0039 100% $2.151.31072M tools
5 DeepSeek V4 Pro 0813
deepseek/deepseek-v4-pro-0813
DeepSeek 0.0039 100% $1.571.048576M tools
6 Grok 4.6
x-ai/grok-4.6
xAI 0.0039 100% $3.00500K tools
7 Kimi K3
moonshotai/kimi-k3
Moonshot AI 0.0039 100% $6.001.048576M tools
8 GPT-5.6 Luna
openai/gpt-5.6-luna
OpenAI 0.0039 100% $0.451.05M tools
9 GPT-5.6 Terra
openai/gpt-5.6-terra
OpenAI 0.0039 100% $4.501.05M tools
10 GPT-5.6 Sol
openai/gpt-5.6-sol
OpenAI 0.0039 100% $4.001.05M tools
11 Grok 4.5
x-ai/grok-4.5
xAI 0.0039 100% $3.00500K tools
12 GLM 5.2
z-ai/glm-5.2
Z.ai (Zhipu) 0.0039 100% $1.491.048576M tools
13 Claude Fable 5
anthropic/claude-fable-5
Anthropic 0.0039 100% $20.001M tools
14 DeepSeek V4 Pro 0423
deepseek/deepseek-v4-pro
DeepSeek 0.0039 100% $1.201.048576M tools
15 Claude Opus 5
anthropic/claude-opus-5
Anthropic 0.0037 95% $10.001M tools
16 DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731
DeepSeek 0.0036 92% $0.101.31072M tools
17 DeepSeek V4 Flash 0423
deepseek/deepseek-v4-flash
DeepSeek 0.0036 92% $0.111.048576M tools
18 Gemini 3.6 Flash
google/gemini-3.6-flash
Google 0.0033 85% $1.501.048576M tools
19 Gemini 3.8 Flash
google/gemini-3.8-flash
Google 0.0031 79% $1.501.048576M tools
20 Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Anthropic 0.0031 79% $6.001M tools
21 Claude Sonnet 5
anthropic/claude-sonnet-5
Anthropic 0.0023 59% $4.001M tools
22 MiniMax M2.7
minimax/minimax-m2.7
MiniMax 0.0021 54% $0.53205K tools
23 MiniMax M3
minimax/minimax-m3
MiniMax 0.0003 8% $0.531.048576M tools
24 Qwen3.7 Max
qwen/qwen3.7-max
Qwen (Alibaba) -0.0015 -38% $2.211M tools
25 Claude Opus 4.8
anthropic/claude-opus-4.8
Anthropic -0.0040 -103% $10.001M tools

Capability: Arena leaderboard dataset (CC-BY-4.0), Arena agent_tool_hallucination, published 2026-09-08. Pricing: OpenRouter, retrieved 2026-09-10. (input x 3 + output x 1) / 4.

What this measures: Human pairwise preference votes. Measures perceived answer quality on this category, not correctness on a fixed test set. For each model this is its score as a share of the board leader's score — 1.0 means parity.

Get this as JSON

curl -s https://agentleaderboards.com/api/v1/route/agent-tool-reliability.json | jq .picks

No key, open CORS, refreshed daily. Set your own budget and constraints →

Other tasks