Clinical Benchmarks

Questions

Short answers about the 13 boards and how to read them.

Which model is best for clinical work?
It depends on the job. Pick the board that matches it, such as documentation, coding or EHR agents, and read that board with the configuration line under each row. The index gives one ordering across boards, as this site's own summary, and the compare page sets two to four models side by side.
How is the index different from a board?
A board is a benchmark's own results on its own scale. The index is this site's calculation: each board's results are placed between that board's lowest and highest model, averaged over the boards a model has, then multiplied down if it has fewer than 3. It is a summary; the boards are the evidence. How it is computed.
How does the index treat a model with results on one or two boards?
Every model with a result on an index board is ranked, but a mean over fewer boards carries less evidence, so it is multiplied down: × 0.577 for one board, × 0.816 for two boards, full weight from three. Gemini 3 Pro scores 84.2 × 0.816 = 68.7 from 2 of 10 boards. The pips under each entry show how many of the index boards its score rests on, and the index views can be filtered to models with results on 2, 3 or 5 or more boards. From three boards up the score is the plain mean.
What do the bar colours mean?
The colour is the lab that built the model, and the lab's mark sits beside the name. A pale bar with an outline is a number the model's own lab reported; a solid bar comes from the benchmark's leaderboard or an independent run.
Why does the same model appear twice on one board?
Because it was run twice under different settings, for example a different reasoning effort, harness or grader. Each row keeps its configuration, and neither is dropped.
Why is a model's rank here different from the official board?
This site may carry fewer rows than the official board. Where the full field size is known, the board page says how many of those models appear here.
Can I use the data?
Yes. The compilation is CC BY 4.0: credit Clinical Benchmarks and link back. CSV and JSON files are on the data page.
How do I report a wrong number?
Use the contact link and name the document and page that shows the right value. Corrections go through the same check as any new row.
Where are the consumer health benchmarks?
Boards about AI people use for their own health are on the sister site linked at the foot of every page.