How rarely the model hallucinates or misuses tool calls during agent runs. Higher is better. Ranked by Arena agent_tool_hallucination, priced against live API rates. 29 models have a published score on this board.
Capability published 2026-09-08 · pricing retrieved 2026-09-10
openai/gpt-6-astradeepseek/deepseek-v4-flash-0731deepseek/deepseek-v4-flash-0731deepseek/deepseek-v4-flash-0731Rows marked frontier are Pareto-optimal — nothing we track is both stronger and cheaper. "Strength" rescales the board so 100% = parity with the leader.
| # | Model | Provider | Score | Strength | $/1M | Context | Tools |
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astraopenai/gpt-6-astra | OpenAI | 0.0039 | 100% | $20.00 | 1.05M | tools |
| 2 | Claude Fable 5.1anthropic/claude-fable-5.1 | Anthropic | 0.0039 | 100% | $20.00 | 1M | tools |
| 3 | GLM 5.3 Flashz-ai/glm-5.3-flash | Z.ai (Zhipu) | 0.0039 | 100% | $0.24 | 1.31072M | tools |
| 4 | GLM 5.3z-ai/glm-5.3 | Z.ai (Zhipu) | 0.0039 | 100% | $2.15 | 1.31072M | tools |
| 5 | DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | DeepSeek | 0.0039 | 100% | $1.57 | 1.048576M | tools |
| 6 | Grok 4.6x-ai/grok-4.6 | xAI | 0.0039 | 100% | $3.00 | 500K | tools |
| 7 | Kimi K3moonshotai/kimi-k3 | Moonshot AI | 0.0039 | 100% | $6.00 | 1.048576M | tools |
| 8 | GPT-5.6 Lunaopenai/gpt-5.6-luna | OpenAI | 0.0039 | 100% | $0.45 | 1.05M | tools |
| 9 | GPT-5.6 Terraopenai/gpt-5.6-terra | OpenAI | 0.0039 | 100% | $4.50 | 1.05M | tools |
| 10 | GPT-5.6 Solopenai/gpt-5.6-sol | OpenAI | 0.0039 | 100% | $4.00 | 1.05M | tools |
| 11 | Grok 4.5x-ai/grok-4.5 | xAI | 0.0039 | 100% | $3.00 | 500K | tools |
| 12 | GLM 5.2z-ai/glm-5.2 | Z.ai (Zhipu) | 0.0039 | 100% | $1.49 | 1.048576M | tools |
| 13 | Claude Fable 5anthropic/claude-fable-5 | Anthropic | 0.0039 | 100% | $20.00 | 1M | tools |
| 14 | DeepSeek V4 Pro 0423deepseek/deepseek-v4-pro | DeepSeek | 0.0039 | 100% | $1.20 | 1.048576M | tools |
| 15 | Claude Opus 5anthropic/claude-opus-5 | Anthropic | 0.0037 | 95% | $10.00 | 1M | tools |
| 16 | DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | DeepSeek | 0.0036 | 92% | $0.10 | 1.31072M | tools |
| 17 | DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | DeepSeek | 0.0036 | 92% | $0.11 | 1.048576M | tools |
| 18 | Gemini 3.6 Flashgoogle/gemini-3.6-flash | 0.0033 | 85% | $1.50 | 1.048576M | tools | |
| 19 | Gemini 3.8 Flashgoogle/gemini-3.8-flash | 0.0031 | 79% | $1.50 | 1.048576M | tools | |
| 20 | Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Anthropic | 0.0031 | 79% | $6.00 | 1M | tools |
| 21 | Claude Sonnet 5anthropic/claude-sonnet-5 | Anthropic | 0.0023 | 59% | $4.00 | 1M | tools |
| 22 | MiniMax M2.7minimax/minimax-m2.7 | MiniMax | 0.0021 | 54% | $0.53 | 205K | tools |
| 23 | MiniMax M3minimax/minimax-m3 | MiniMax | 0.0003 | 8% | $0.53 | 1.048576M | tools |
| 24 | Qwen3.7 Maxqwen/qwen3.7-max | Qwen (Alibaba) | -0.0015 | -38% | $2.21 | 1M | tools |
| 25 | Claude Opus 4.8anthropic/claude-opus-4.8 | Anthropic | -0.0040 | -103% | $10.00 | 1M | tools |
Capability: Arena leaderboard dataset (CC-BY-4.0), Arena agent_tool_hallucination, published 2026-09-08. Pricing: OpenRouter, retrieved 2026-09-10. (input x 3 + output x 1) / 4.
curl -s https://agentleaderboards.com/api/v1/route/agent-tool-reliability.json | jq .picks
No key, open CORS, refreshed daily. Set your own budget and constraints →
Writing & editing code · Autonomous agent work · Following operator instructions · Recovering from shell errors · Hard reasoning · Mathematics · Precise instruction following · Long multi-turn chat · Long inputs · Creative writing · General assistant · Software & IT services · Legal & government · Medicine & healthcare · Science · Business & finance · Writing & language