Best AI model for medicine & healthcare

Clinical and healthcare prompts. Not medical advice; check outputs. Ranked by Arena text / industry_medicine_and_healthcare, priced against live API rates. 44 models have a published score on this board.

Capability published 2026-09-02 · pricing retrieved 2026-09-10

Most capable
Claude Fable 5.1
Anthropic
$20.00 / 1M blended
100% of leader strength
anthropic/claude-fable-5.1
Best value
Gemini 3.8 Flash
Google
$1.50 / 1M blended
93% of leader strength · 93% cheaper
google/gemini-3.8-flash
Budget
DeepSeek V4 Flash 0731
DeepSeek
$0.10 / 1M blended
80% of leader strength · 100% cheaper
deepseek/deepseek-v4-flash-0731
Best for agents
Gemini 3.8 Flash
Google
$1.50 / 1M blended
93% of leader strength · 93% cheaper
google/gemini-3.8-flash

Top 25 models for medicine & healthcare

Rows marked frontier are Pareto-optimal — nothing we track is both stronger and cheaper. "Strength" rescales the board so 100% = parity with the leader.

#ModelProviderScoreStrength$/1MContextTools
1 Claude Fable 5.1
anthropic/claude-fable-5.1
Anthropic 1518 100% $20.001M tools
2 Claude Opus 5
anthropic/claude-opus-5
Anthropic 1517 100% $10.001M tools
3 Qwen3.8 Max (0902)
qwen/qwen3.8-max-0902
Qwen (Alibaba) 1510 97% $3.001M tools
4 Claude Opus 4.6
anthropic/claude-opus-4.6
Anthropic 1501 95% $10.001M tools
5 Gemini 3.8 Flash
google/gemini-3.8-flash
Google 1495 93% $1.501.048576M tools
6 Gemini 3.5 Flash
google/gemini-3.5-flash
Google 1492 92% $3.381.048576M tools
7 Claude Opus 4.7
anthropic/claude-opus-4.7
Anthropic 1490 92% $10.001M tools
8 GLM 5.3
z-ai/glm-5.3
Z.ai (Zhipu) 1489 92% $2.151.31072M tools
9 Kimi K3
moonshotai/kimi-k3
Moonshot AI 1484 90% $6.001.048576M tools
10 Gemini 3.7 Flash
google/gemini-3.7-flash
Google 1483 90% $1.501.048576M tools
11 Claude Fable 5
anthropic/claude-fable-5
Anthropic 1476 88% $20.001M tools
12 GLM 5.1
z-ai/glm-5.1
Z.ai (Zhipu) 1472 87% $1.49205K tools
13 GLM 5.2
z-ai/glm-5.2
Z.ai (Zhipu) 1471 87% $1.491.048576M tools
14 Qwen3.7 Max
qwen/qwen3.7-max
Qwen (Alibaba) 1470 86% $2.211M tools
15 Gemini 3.6 Flash
google/gemini-3.6-flash
Google 1469 86% $1.501.048576M tools
16 Qwen3.7 Plus
qwen/qwen3.7-plus
Qwen (Alibaba) 1465 85% $0.561M tools
17 DeepSeek V4 Pro 0813
deepseek/deepseek-v4-pro-0813
DeepSeek 1465 85% $1.571.048576M tools
18 DeepSeek V4 Pro 0423
deepseek/deepseek-v4-pro
DeepSeek 1465 85% $1.201.048576M tools
19 GLM 5.3 Flash
z-ai/glm-5.3-flash
Z.ai (Zhipu) 1459 83% $0.241.31072M tools
20 Claude Opus 4.8
anthropic/claude-opus-4.8
Anthropic 1458 83% $10.001M tools
21 Kimi K2.6
moonshotai/kimi-k2.6
Moonshot AI 1454 82% $1.71262K tools
22 GLM 5
z-ai/glm-5
Z.ai (Zhipu) 1454 82% $0.93205K tools
23 Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Anthropic 1454 82% $6.001M tools
24 Kimi K2.5
moonshotai/kimi-k2.5
Moonshot AI 1449 80% $0.90262K tools
25 DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731
DeepSeek 1449 80% $0.101.31072M tools

Capability: Arena leaderboard dataset (CC-BY-4.0), Arena text / industry_medicine_and_healthcare, published 2026-09-02. Pricing: OpenRouter, retrieved 2026-09-10. (input x 3 + output x 1) / 4.

What this measures: Human pairwise preference votes. Measures perceived answer quality on this category, not correctness on a fixed test set. For each model this is the probability it beats the board leader in a head-to-head human vote — 0.50 means a dead tie and is the best possible value.

Get this as JSON

curl -s https://agentleaderboards.com/api/v1/route/medicine.json | jq .picks

No key, open CORS, refreshed daily. Set your own budget and constraints →

Other tasks