Clinical Benchmarks

Method

Where each number comes from, how it is checked, and how the index is built.

Where numbers come from

Every result traces to one of three kinds of document. Bar colour on the charts is the lab; a pale bar with an outline is a number the model's own lab reported.

  • Official leaderboard. The benchmark's owner ran the model and published the score on their own board or in their own paper.
  • Independent run. A party that neither built the model nor owns the benchmark ran it and published the result with its configuration.
  • Vendor-reported. The lab that built the model printed the score in its system card, model card or launch post. The lab chose the settings, and sometimes the grader.

Right now the 425 results split as 332 official leaderboard, 211 independent run and 50 vendor-reported. Each board chart can be filtered by who reported the number.

How a row is checked

A row is published only when all four of these are recorded:

  • the document, preferring the first-party original over any page that repeats it;
  • the number exactly as the document prints it, with its own rounding and its own percent sign or lack of one;
  • a verbatim quote from the document that contains the number and shows which model and which measure it belongs to;
  • a locator: page and table, section, figure, or the board's row and the date the board showed it.

The configuration the source states goes into the row as well: reasoning effort, harness, grader, whether a length adjustment was applied. It appears under the model name on every chart and in full in the row's sources. Two rows for the same model on one board almost always differ in configuration, and both are kept.

Printed form and chart form can differ. A card that prints 66.0 on a board scored 0 to 1 is charted at 0.660; the row's sources show both. Nothing is rounded beyond what the source printed.

When documents disagree

A first-party document outranks a page that mirrors it, and a benchmark owner's board outranks a secondary write-up of that board. The row keeps the number from its document of record.

Other documents that print the same result are listed in the row's sources, each with the number it printed: “also printed in” when it matches, “a different number in” when it does not. When the disagreement cannot be settled from the documents, the row carries the disputed marker.

Reading a board

Rank means position among the rows on that board and that measure only. Some boards list many more models than this site carries; where the board's full field size is known, pages say so as “N of M models on the board”. A rank here can be higher than the model's rank on the official board for that reason.

Vendor-reported rows on the same board may use different graders or settings from each other. The configuration line under each row is there so a reader can decide whether two rows are comparable for their purpose.

The index

The Clinical Benchmarks Index is this site's own calculation, made from the board results and nothing else. It is a summary for readers who want one ordering; it is not a benchmark, and the boards remain the evidence.

  • Per board. A model's best headline result, placed between the lowest and highest such result on that board, on a 0 to 100 scale.
  • Across boards. The index is the mean over the boards the model has results on, times a confidence multiplier.
  • Confidence multiplier. A mean over fewer boards rests on less evidence, so it counts for less: full from 3 boards; below that the square root of boards / 3 (2 boards x0.816, 1 board x0.577). The multiplier is about how many boards a score rests on; it is separate from the confidence markers on single rows.
  • Who is ranked. Every model with a result on at least one index board. Left out: harness plus model pairings, products, research systems, human and baseline references.
  • Which boards. A board needs at least 5 such models. Every entry shows how many boards it covers and its result on each.
  • Coverage. The index views mark each entry's coverage with pips and can be filtered to models with results on 2, 3 or 5 or more boards. A filter hides entries; it does not change any score or rank.

Every index score is shown with its working. Gemini 3 Pro has a mean placement of 84.2 over 2 of 10 boards, so its index is 84.2 × 0.816 = 68.7. The printed figures multiply out to the score. From three boards up the score is the plain mean. Ranks follow the score after the multiplier.

Boards and their range in the index
BoardModelsLowestHighestIn the index
HealthBench Professional260.350.703Yes
MedScribe (Vals AI)944.2791.43Yes
MedCode (Vals AI)9419.7263.57Yes
MedPIC283680.7Yes
MedXpertQA (MM)2223.581.5Yes
MAST (Medical AI Superintelligence Test)853.760.2Yes
MedHELM100.3420.652Yes
First, Do NOHARM (v2)1355.879.7Yes
PhysicianBench175.368.4Yes
EHR-Complex170.160.65Yes
HealthAgentBench0––No, fewer than 5 models
CHI-Bench0––No, fewer than 5 models
HealthAdminBench3––No, fewer than 5 models

148 models are ranked: 55 from one board, 42 from three or more.

Confidence markers

Rows that have been matched against their document carry no marker. The other two states are always visible, on charts, tables and model pages:

  • Unverified the number has been recorded but not yet matched against its source document.
  • Disputed documents disagree about the number; the row's sources show each one.

What we never do

  • Combine, average or rank scores from different benchmarks anywhere except the index, which is labelled as this site's own calculation wherever it appears. Each board ranks only its own rows.
  • Add precision a source did not print, or convert a score to another scale and present it as the source's.
  • Publish a number without its document. A row whose document is not yet recorded shows “source pending”.

What gets a board

A benchmark is tracked when it tests clinical work (reasoning, documentation, coding, EHR tasks, administrative work or patient safety) and has recent results for current models. Benchmarks that saturated or stopped reporting current models are listed with the reason at the foot of the boards page. Benchmarks about AI that people use for their own health are covered by the sister site linked in the footer.