Clinical Benchmarks

Boards by clinical job

Each board tests one clinical job. The strip shows every row on the board, with the leader marked.

Clinical reasoning and knowledge

Consults, diagnosis and management questions, image-based cases and broad clinical suites.

  • HealthBench Professional

    Physician-selected workplace tasks, from care consults to documentation and research, graded on physician-written rubrics.

    OpenAI, 525 tasks

    Leads this board
    31 rowsmeasured to Sep 2026
  • MedXpertQA (MM)

    Expert-level multiple-choice questions over clinical images across 17 specialties.

    Tsinghua University, 2,000 multimodal questions

    Leads this board
    22 rows of 5 on the boardmeasured to Aug 2026
  • MAST (Medical AI Superintelligence Test)

    A composite of clinical benchmarks spanning diagnostic and management reasoning, safety, multimodal imaging, and agentic capability.

    ARISE AI Research Network, 6-benchmark composite

    Leads this board
    8 rows of 11 on the boardmeasured to Aug 2026
  • MedHELM

    Holistic clinical evaluation across 121 tasks in a clinician-validated taxonomy, ranked by mean win rate.

    Stanford CRFM, 121 tasks

    Leads this board
    10 rows of 11 on the boardmeasured to May 2026

Documentation and coding

Notes written from visits and diagnosis codes assigned to hospital stays.

  • MedScribe (Vals AI)

    Quality and compliance of SOAP notes generated from clinical visits, scored against documentation rubrics.

    Vals AI, 100 SOAP-note cases

    Leads this board
    105 rowsmeasured to Sep 2026
  • MedCode (Vals AI)

    ICD-10-CM coding of whole hospital stays from discharge summaries and notes, checked against professional coders.

    Vals AI, 2,755 patient records

    Leads this board
    103 rowsmeasured to Sep 2026

EHR and workflow agents

Agents that query records, run multi-step physician workflows and complete administrative work.

  • PhysicianBench

    Agents carrying out long-horizon physician workflows inside real EHR systems, verified by execution against those systems.

    academic team, 100 clinical tasks

    Leads this board
    21 rows of 12 on the boardmeasured to Sep 2026
  • EHR-Complex

    Agentic clinical reasoning over MIMIC-IV records through SQL and Python, at patient and population level.

    academic team, 3,915-task test set

    Leads this board
    18 rowsmeasured to Jun 2026
  • HealthAgentBench

    Agent harnesses completing realistic terminal-based healthcare tasks built from real clinical artifacts.

    Microsoft Research, 54 agentic tasks

    Leads this board
    12 rowsmeasured to Jul 2026
  • CHI-Bench

    Long-horizon healthcare operations workflows for agents: prior authorization, utilization management, and care management.

    actAVA, 75 operations workflows

    Leads this board
    43 rows of 45 on the boardmeasured to Aug 2026
  • HealthAdminBench

    Computer-use agents completing healthcare administration workflows: prior authorizations, denial appeals, and DME ordering.

    Kinetic Systems, 135 admin tasks

    7 rows of 5 on the boardmeasured to Apr 2026

Safety

Whether recommendations could harm a patient, and whether safety rules hold when patient details change.

  • MedPIC

    Tests whether models apply and withdraw medication-safety rules correctly as patient information changes.

    Hou et al., %

    Leads this board
    28 rowsmeasured to Aug 2026
  • First, Do NOHARM (v2)

    How often, and how severely, model consultation recommendations contain potentially harmful errors.

    Stanford/Harvard consortium, 1,100 consultation cases

    Leads this board
    LiSA 2.586.2
    17 rows of 19 on the boardmeasured to Sep 2026

Benchmarks not tracked

Considered and left out, with the reason. As of 30 September 2026.

  • HealthBench ConsensusNear-saturated physician-consensus baseline; frontier runs stopped reporting it separately.
  • MedQA / MultiMedQAExam-style multiple choice, saturated above 95 percent since 2025; archived by its trackers.
  • AgentClinicNo public frontier-model results since 2025.
  • CRAFT-MDNo public frontier-model results since 2025.
  • MedAgentBenchV2 lives on inside the MAST composite; the standalone board has no current frontier rows.
  • SDBench / MAI-DxOMicrosoft's 2025 sequential-diagnosis study was not re-run on current models.
  • Open Medical-LLM Leaderboard (Hugging Face)Built on saturated exam sets; no frontier submissions in 2026.
  • MedArenaClinician preference arena; ratings pool too thin on current frontier models to quote.
  • AMIE evaluationsGoogle DeepMind research prototypes, never opened to cross-vendor comparison.
  • LiveClin, PrIME-LLM, MedMCP-CalcSingle studies with two or fewer current-frontier rows; tracked for a future qualifying update.