Clinical Benchmarks Index
Each model’s results placed between the lowest and highest model on each of 10 clinical boards, 0 to 100, averaged over the boards it has, then scaled down if it has fewer than 3. This site’s own calculation.
- Models ranked
- 148
- Index boards
- 10
- On 3+ boards
- 42
- Updated
- 30 Sep 2026
Latest changes
All updates- CHI-Benchopenai-agents + Nemotron 3 Ultra 256K, openclaw + grok-4.3, deepagents + grok-4.3 and others30 Sep 202631 added
- 93 added
Clinical Benchmarks Index
This site’s own calculation from 10 boards, on a 0 to 100 scale. How it is computed
- Anthropic
- OpenAI
- Alibaba
- Meta
- Moonshot AI
- Other labs
- Index boards behind the score
- 20 of 148 ranked models
Every model with a result on an index board is ranked. Its score is the mean of its board placements times a confidence multiplier: × 0.577 for one board, × 0.816 for two boards, full weight from three. The pips show how many of the 10 index boards a score rests on. Filtering keeps each model’s rank among all 148.
Leader on each index board
Each on its own scale- HealthBench Professional0.703GPT-6 Astra (Anthropic run)Anthropic public-API reproduction0.703Anthropic public-API reproduction
- MedScribe (Vals AI)91.43Claude Opus 5.5model ID anthropic/claude-opus-5-591.43model ID anthropic/claude-opus-5-5
- MedCode (Vals AI)63.57Claude Opus 563.57
- MedPIC80.7Gemini-3.1-Prozero-shot80.7zero-shot
- MedXpertQA (MM)81.5GPT-5.6 SolQwen-run comparison in the Qwen381.5Qwen-run comparison in the Qwen3
- MAST (Medical AI Superintelligence Test)60.2GPT-5.6 SolMAST in preview60.2MAST in preview
- MedHELM0.652Gemini 3.1 Pro (Preview)0.652
- First, Do NOHARM (v2)86.2LiSA 2.5First Do NOHARM v2 overall86.2First Do NOHARM v2 overall
- PhysicianBench68.4Claude Opus 5.5 (max)Anthropic harness68.4Anthropic harness
- EHR-Complex0.650GPT-5.4 (high reasoning)average over 12 intent columns0.650average over 12 intent columns
Index and API price
Blended list price per million tokens, 3 parts input to 1 part output. Log scale.
77 of 148 models
The line joins models that no cheaper model beats on the index.
Index and release date
Each model at its public release date.
116 of 148 models
The line is the highest index reached by that date.
Lab standings
Each lab’s highest model in the index, with the boards behind its score.
- AnthropicClaude Sonnet 5.5, rank 1, 16 ranked4 of 10 index boards91.0
- MetaMuse Spark 1.1, rank 3, 7 ranked3 of 10 index boards88.6
- GoogleGemini 3.5 Flash, rank 4, 26 ranked3 of 10 index boards87.4
- OpenAIGPT-6 Astra, rank 5, 24 ranked3 of 10 index boards87.2
- Moonshot AIKimi K3, rank 7, 3 ranked4 of 10 index boards84.3
- AlibabaQwen3.8 Max, rank 13, 20 ranked3 of 10 index boards79.5
- SpaceX AIGrok 4.7, rank 16, 10 ranked3 of 10 index boards75.7
- MiniMaxMiniMax M3, rank 35, 3 ranked2 of 10 index boards63.6
- XiaomiMiMo V2.6 Pro, rank 37, 4 ranked2 of 10 index boards62.8
- ZhipuGLM 5.3, rank 41, 5 ranked2 of 10 index boards61.1
All 25 labsFewer labs
- DeepSeekDeepSeek V4 Pro, rank 44, 7 ranked4 of 10 index boards60.0
- TencentHy4 Preview, rank 462 of 10 index boards59.1
- Thinking MachinesInkling, rank 49, 2 ranked2 of 10 index boards58.0
- Ant GroupLing 3.0 Flash, rank 68, 4 ranked2 of 10 index boards47.6
- MistralMistral Medium 3.5, rank 862 of 10 index boards42.8
- InceptionMercury 2.5, rank 1062 of 10 index boards34.6
- PoolsideLaguna M.1, rank 110, 2 ranked2 of 10 index boards32.0
- UbiquantFleming-R1-32B, rank 111, 2 ranked1 of 10 index boards31.7
- NVIDIANemotron 3 Ultra, rank 121, 2 ranked1 of 10 index boards24.9
- FreedomIntelligenceHuatuoGPT-o1-70B, rank 122, 2 ranked1 of 10 index boards24.8
- CohereCommand A+, rank 1232 of 10 index boards24.1
- BaichuanBaichuan-M2-32B, rank 1321 of 10 index boards15.2
- ZJU4HealthCareHealthGPT-Pro-8B, rank 135, 2 ranked1 of 10 index boards8.5
- ZJU-AI4HHulu-Med-7B, rank 1361 of 10 index boards7.7
- MicrosoftMAI-Thinking-1, rank 1461 of 10 index boards0.0
Earlier changes
All updates- 92 added
- HealthBench ProfessionalClaude Opus 4.8 (Opus 4.8 grader), Claude Fable 5 (September card), GPT-6 Astra and others30 Sep 20269 added, 1 rescored
- 12 added
- 12 added
- 7 added
- 1 added
- 5 added
- 28 added
Boards
All boards with details- Clinical reasoning and knowledgeHealthBench Professional
Physician-selected workplace tasks, from care consults to documentation and research, graded on physician-written rubrics.
- 1GPT-6 Astra (Anthropic run)0.703
- 2Claude Sonnet 5.50.692
- 3Claude Fable 50.660
- 4Claude Opus 5.50.656
- 5GPT-6 Astra0.647
31 rows, to Sep 2026Full board - Clinical reasoning and knowledgeMedXpertQA (MM)
Expert-level multiple-choice questions over clinical images across 17 specialties.
- 1GPT-5.6 Sol81.5
- 2Gemini 3.1 Pro81.3
- 3Qwen3.8 Max80.4
- 4Claude Fable 580.0
- 5Muse Spark78.4
22 rows, to Aug 2026Full board - Clinical reasoning and knowledgeMAST (Medical AI Superintelligence Test)
A composite of clinical benchmarks spanning diagnostic and management reasoning, safety, multimodal imaging, and agentic capability.
- 1GPT-5.6 Sol60.2
- 2Kimi K360.1
- 3Gemini 3.6 Flash59.3
- 4Gemini 3.1 Pro58.9
- 5Qwen3.5 397B A17B57.9
8 rows, to Aug 2026Full board - Clinical reasoning and knowledgeMedHELM
Holistic clinical evaluation across 121 tasks in a clinician-validated taxonomy, ranked by mean win rate.
- 1Gemini 3.1 Pro (Preview)0.652
- 2Gemini 3.5 Flash0.642
- 3Muse Spark (2026-04-08)0.621
- 4GPT-5.4 mini0.552
- 5GPT-5.4 (2026-03-05)0.538
10 rows, to May 2026Full board - Documentation and codingMedScribe (Vals AI)
Quality and compliance of SOAP notes generated from clinical visits, scored against documentation rubrics.
- 1Claude Opus 5.591.43
- 2Claude Fable 5.191.29
- 3Claude Sonnet 5.591.10
- 4Claude Opus 590.98
- 5Muse Spark 1.290.06
105 rows, to Sep 2026Full board - Documentation and codingMedCode (Vals AI)
ICD-10-CM coding of whole hospital stays from discharge summaries and notes, checked against professional coders.
- 1Claude Opus 563.57
- 2Gemini 3.1 Pro Preview (02/26)59.06
- 3Claude Fable 556.07
- 4Gemini 3 Flash (12/25)55.92
- 5Gemini 3.5 Flash55.83
103 rows, to Sep 2026Full board - EHR and workflow agentsPhysicianBench
Agents carrying out long-horizon physician workflows inside real EHR systems, verified by execution against those systems.
- 1Claude Opus 5.5 (max)68.4
- 2Claude Sonnet 5.5 (max)63.2
- 3Claude Fable 5.1 (max)61.0
- 4Claude Opus 5 (max)57.6
- 5Claude Sonnet 5.5 (xhigh)56.4
21 rows, to Sep 2026Full board - EHR and workflow agentsEHR-Complex
Agentic clinical reasoning over MIMIC-IV records through SQL and Python, at patient and population level.
- 1GPT-5.4 (high reasoning)0.650
- 2Gemini 3.1 Pro0.630
- 3Kimi-K2.50.620
- 3Qwen3.5-397B0.620
- 5DeepSeek-V3.2-Exp0.590
18 rows, to Jun 2026Full board - EHR and workflow agentsHealthAgentBench
Agent harnesses completing realistic terminal-based healthcare tasks built from real clinical artifacts.
- 1Claude Code (Opus 5)55
- 2Codex (GPT-5.6-sol)45
- 3Codex (GPT 5.5)42
- 4Copilot (Opus 4.8)36
- 5Copilot (GPT 5.5)35
12 rows, to Jul 2026Full board - EHR and workflow agentsCHI-Bench
Long-horizon healthcare operations workflows for agents: prior authorization, utilization management, and care management.
- 1erius + claude-opus-554.7
- 2claude-code + claude-opus-537.3
- 2erius + claude-opus-4-837.3
- 4claude-code + claude-opus-4-833.3
- 5claude-code + claude-opus-4-628.0
43 rows, to Aug 2026Full board - EHR and workflow agentsHealthAdminBench
Computer-use agents completing healthcare administration workflows: prior authorizations, denial appeals, and DME ordering.
- 1Claude Opus 4.6 (computer-use agent)36.3
- 2GPT-5.4 (computer-use agent)26.7
- 3Kimi K2.515.6
- 4Claude Opus 4.6 (standardized harness)14.8
- 5Qwen 3.513.3
7 rows, to Apr 2026Full board - SafetyMedPIC
Tests whether models apply and withdraw medication-safety rules correctly as patient information changes.
- 1Gemini-3.1-Pro80.7
- 2GPT-575.6
- 3Qwen3.5-Plus72.4
- 4DeepSeek-V4-Pro71.7
- 5GPT-OSS-120B70.7
28 rows, to Aug 2026Full board - SafetyFirst, Do NOHARM (v2)
How often, and how severely, model consultation recommendations contain potentially harmful errors.
- 1LiSA 2.586.2
- 2Doximity Ask 6.184.5
- 3OpenEvidence80.0
- 4Glass 5.6 Max79.7
- 4Muse Spark 1.179.7
17 rows, to Sep 2026Full board