Clinical Benchmarks

Clinical Benchmarks Index

Each model’s results placed between the lowest and highest model on each of 10 clinical boards, 0 to 100, averaged over the boards it has, then scaled down if it has fewer than 3. This site’s own calculation.

Models ranked
148
Index boards
10
On 3+ boards
42
Updated
30 Sep 2026

See the full indexHow it is computed

Latest changes

All updates
  1. CHI-Bench
    openai-agents + Nemotron 3 Ultra 256K, openclaw + grok-4.3, deepagents + grok-4.3 and others
    30 Sep 2026
    31 added
  2. MedScribe (Vals AI)
    Nemotron 3.5 Lightning, Llama 4 Scout, Llama 4 Maverick and others
    30 Sep 2026
    93 added

Clinical Benchmarks Index

This site’s own calculation from 10 boards, on a 0 to 100 scale. How it is computed

Boards behind each score
  • Anthropic
  • Google
  • OpenAI
  • Alibaba
  • Meta
  • Moonshot AI
  • Other labs
  • Index boards behind the score
  • 20 of 148 ranked models

Every model with a result on an index board is ranked. Its score is the mean of its board placements times a confidence multiplier: × 0.577 for one board, × 0.816 for two boards, full weight from three. The pips show how many of the 10 index boards a score rests on. Filtering keeps each model’s rank among all 148.

Leader on each index board

Each on its own scale

Index and API price

Blended list price per million tokens, 3 parts input to 1 part output. Log scale.

77 of 148 models

020406080100$0.01$0.03$0.1$0.3$1$3$10$30$100

The line joins models that no cheaper model beats on the index.

Index and release date

Each model at its public release date.

116 of 148 models

020406080100Apr 2024Oct 2024Apr 2025Oct 2025Apr 2026Oct 2026

The line is the highest index reached by that date.

Lab standings

Each lab’s highest model in the index, with the boards behind its score.

  1. AnthropicClaude Sonnet 5.5, rank 1, 16 ranked4 of 10 index boards91.0
  2. MetaMuse Spark 1.1, rank 3, 7 ranked3 of 10 index boards88.6
  3. GoogleGemini 3.5 Flash, rank 4, 26 ranked3 of 10 index boards87.4
  4. OpenAIGPT-6 Astra, rank 5, 24 ranked3 of 10 index boards87.2
  5. Moonshot AIKimi K3, rank 7, 3 ranked4 of 10 index boards84.3
  6. AlibabaQwen3.8 Max, rank 13, 20 ranked3 of 10 index boards79.5
  7. SpaceX AIGrok 4.7, rank 16, 10 ranked3 of 10 index boards75.7
  8. MiniMaxMiniMax M3, rank 35, 3 ranked2 of 10 index boards63.6
  9. XiaomiMiMo V2.6 Pro, rank 37, 4 ranked2 of 10 index boards62.8
  10. ZhipuGLM 5.3, rank 41, 5 ranked2 of 10 index boards61.1
All 25 labsFewer labs
  1. DeepSeekDeepSeek V4 Pro, rank 44, 7 ranked4 of 10 index boards60.0
  2. TencentHy4 Preview, rank 462 of 10 index boards59.1
  3. Thinking MachinesInkling, rank 49, 2 ranked2 of 10 index boards58.0
  4. Ant GroupLing 3.0 Flash, rank 68, 4 ranked2 of 10 index boards47.6
  5. MistralMistral Medium 3.5, rank 862 of 10 index boards42.8
  6. InceptionMercury 2.5, rank 1062 of 10 index boards34.6
  7. PoolsideLaguna M.1, rank 110, 2 ranked2 of 10 index boards32.0
  8. UbiquantFleming-R1-32B, rank 111, 2 ranked1 of 10 index boards31.7
  9. NVIDIANemotron 3 Ultra, rank 121, 2 ranked1 of 10 index boards24.9
  10. FreedomIntelligenceHuatuoGPT-o1-70B, rank 122, 2 ranked1 of 10 index boards24.8
  11. CohereCommand A+, rank 1232 of 10 index boards24.1
  12. BaichuanBaichuan-M2-32B, rank 1321 of 10 index boards15.2
  13. ZJU4HealthCareHealthGPT-Pro-8B, rank 135, 2 ranked1 of 10 index boards8.5
  14. ZJU-AI4HHulu-Med-7B, rank 1361 of 10 index boards7.7
  15. MicrosoftMAI-Thinking-1, rank 1461 of 10 index boards0.0

Earlier changes

All updates
  1. MedCode (Vals AI)
    Command A+, Laguna XS.2, Laguna M.1 and others
    30 Sep 2026
    92 added
  2. HealthBench Professional
    Claude Opus 4.8 (Opus 4.8 grader), Claude Fable 5 (September card), GPT-6 Astra and others
    30 Sep 2026
    9 added, 1 rescored
  3. EHR-Complex
    Qwen3-4B, Qwen3-14B, Gemini 2.5 Pro and others
    30 Sep 2026
    12 added
  4. PhysicianBench
    MiniMax M2.7, MiMo-v2.5-Pro, DeepSeek V4-Pro and others
    30 Sep 2026
    12 added
  5. MedXpertQA (MM)
    Gemma 4 E2B, Gemma 4 E4B, Qwen3-VL-235B-A22B and others
    30 Sep 2026
    7 added
  6. HealthAgentBench
    Codex (GPT 5.3)
    30 Sep 2026
    1 added
  7. First, Do NOHARM (v2)
    GLM 5.1, Glass 5.6 Max, OpenEvidence and others
    30 Sep 2026
    5 added
  8. MedPIC
    Gemini-3.1-Pro, GPT-5, Qwen3.5-Plus and others
    30 Sep 2026
    28 added
  • Clinical reasoning and knowledgeHealthBench Professional

    Physician-selected workplace tasks, from care consults to documentation and research, graded on physician-written rubrics.

    1. 1GPT-6 Astra (Anthropic run)0.703
    2. 2Claude Sonnet 5.50.692
    3. 3Claude Fable 50.660
    4. 4Claude Opus 5.50.656
    5. 5GPT-6 Astra0.647
    31 rows, to Sep 2026Full board
  • Clinical reasoning and knowledgeMedXpertQA (MM)

    Expert-level multiple-choice questions over clinical images across 17 specialties.

    1. 1GPT-5.6 Sol81.5
    2. 2Gemini 3.1 Pro81.3
    3. 3Qwen3.8 Max80.4
    4. 4Claude Fable 580.0
    5. 5Muse Spark78.4
    22 rows, to Aug 2026Full board
  • Clinical reasoning and knowledgeMAST (Medical AI Superintelligence Test)

    A composite of clinical benchmarks spanning diagnostic and management reasoning, safety, multimodal imaging, and agentic capability.

    1. 1GPT-5.6 Sol60.2
    2. 2Kimi K360.1
    3. 3Gemini 3.6 Flash59.3
    4. 4Gemini 3.1 Pro58.9
    5. 5Qwen3.5 397B A17B57.9
    8 rows, to Aug 2026Full board
  • Clinical reasoning and knowledgeMedHELM

    Holistic clinical evaluation across 121 tasks in a clinician-validated taxonomy, ranked by mean win rate.

    1. 1Gemini 3.1 Pro (Preview)0.652
    2. 2Gemini 3.5 Flash0.642
    3. 3Muse Spark (2026-04-08)0.621
    4. 4GPT-5.4 mini0.552
    5. 5GPT-5.4 (2026-03-05)0.538
    10 rows, to May 2026Full board
  • Documentation and codingMedScribe (Vals AI)

    Quality and compliance of SOAP notes generated from clinical visits, scored against documentation rubrics.

    1. 1Claude Opus 5.591.43
    2. 2Claude Fable 5.191.29
    3. 3Claude Sonnet 5.591.10
    4. 4Claude Opus 590.98
    5. 5Muse Spark 1.290.06
    105 rows, to Sep 2026Full board
  • Documentation and codingMedCode (Vals AI)

    ICD-10-CM coding of whole hospital stays from discharge summaries and notes, checked against professional coders.

    1. 1Claude Opus 563.57
    2. 2Gemini 3.1 Pro Preview (02/26)59.06
    3. 3Claude Fable 556.07
    4. 4Gemini 3 Flash (12/25)55.92
    5. 5Gemini 3.5 Flash55.83
    103 rows, to Sep 2026Full board
  • EHR and workflow agentsPhysicianBench

    Agents carrying out long-horizon physician workflows inside real EHR systems, verified by execution against those systems.

    1. 1Claude Opus 5.5 (max)68.4
    2. 2Claude Sonnet 5.5 (max)63.2
    3. 3Claude Fable 5.1 (max)61.0
    4. 4Claude Opus 5 (max)57.6
    5. 5Claude Sonnet 5.5 (xhigh)56.4
    21 rows, to Sep 2026Full board
  • EHR and workflow agentsEHR-Complex

    Agentic clinical reasoning over MIMIC-IV records through SQL and Python, at patient and population level.

    1. 1GPT-5.4 (high reasoning)0.650
    2. 2Gemini 3.1 Pro0.630
    3. 3Kimi-K2.50.620
    4. 3Qwen3.5-397B0.620
    5. 5DeepSeek-V3.2-Exp0.590
    18 rows, to Jun 2026Full board
  • EHR and workflow agentsHealthAgentBench

    Agent harnesses completing realistic terminal-based healthcare tasks built from real clinical artifacts.

    1. 1Claude Code (Opus 5)55
    2. 2Codex (GPT-5.6-sol)45
    3. 3Codex (GPT 5.5)42
    4. 4Copilot (Opus 4.8)36
    5. 5Copilot (GPT 5.5)35
    12 rows, to Jul 2026Full board
  • EHR and workflow agentsCHI-Bench

    Long-horizon healthcare operations workflows for agents: prior authorization, utilization management, and care management.

    1. 1erius + claude-opus-554.7
    2. 2claude-code + claude-opus-537.3
    3. 2erius + claude-opus-4-837.3
    4. 4claude-code + claude-opus-4-833.3
    5. 5claude-code + claude-opus-4-628.0
    43 rows, to Aug 2026Full board
  • EHR and workflow agentsHealthAdminBench

    Computer-use agents completing healthcare administration workflows: prior authorizations, denial appeals, and DME ordering.

    1. 1Claude Opus 4.6 (computer-use agent)36.3
    2. 2GPT-5.4 (computer-use agent)26.7
    3. 3Kimi K2.515.6
    4. 4Claude Opus 4.6 (standardized harness)14.8
    5. 5Qwen 3.513.3
    7 rows, to Apr 2026Full board
  • SafetyMedPIC

    Tests whether models apply and withdraw medication-safety rules correctly as patient information changes.

    1. 1Gemini-3.1-Pro80.7
    2. 2GPT-575.6
    3. 3Qwen3.5-Plus72.4
    4. 4DeepSeek-V4-Pro71.7
    5. 5GPT-OSS-120B70.7
    28 rows, to Aug 2026Full board
  • SafetyFirst, Do NOHARM (v2)

    How often, and how severely, model consultation recommendations contain potentially harmful errors.

    1. 1LiSA 2.586.2
    2. 2Doximity Ask 6.184.5
    3. 3OpenEvidence80.0
    4. 4Glass 5.6 Max79.7
    5. 4Muse Spark 1.179.7
    17 rows, to Sep 2026Full board