About Clinical Benchmarks
Clinical Benchmarks gathers published results of AI models on benchmarks built around clinical work: answering clinicians' questions, writing notes, assigning codes, operating an EHR, doing administrative work, avoiding harmful recommendations. It currently covers 13 benchmarks and 425 results.
It is written for the people who choose models for clinical settings: CMIOs, informatics teams, AI governance committees, and the researchers and evaluation teams who build and test these systems. The pages are laid out to be printed and cited, and every number carries a link to the document it came from.
What it is not
- A ranking of models overall. Each board stands alone.
- A certification. A high score on a benchmark is evidence about that benchmark's tasks, not approval for clinical use.
- A source of new scores. Numbers come from system cards, official leaderboards, papers and independent runs; the method page explains how each is checked.
Sister site
Benchmarks about AI people use for their own health are covered by a sister site, linked at the foot of every page.
Corrections
If a number here does not match its document, get in touch with the page and table.