Clinical Benchmarks

MedHELM: current results

Stanford CRFM / HAI and multi-institution collaborators · 121 tasks / 31 datasets · index updated August 16, 2026

Gemini 3.1 Pro (Preview) holds the top current result on MedHELM, 0.652 as of 2026-05, per MedHELM leaderboard (medhelm.org), v5.0.0. Holistic evaluation of LLMs on 121 clinical tasks across 5 categories and 22 subcategories (31 datasets) in a clinician-validated taxonomy; ranked by mean win rate.

Current results

Result detail

#modelscoreas of
1Google logoGemini 3.1 Pro (Preview) Google0.6522026-05
2Google logoGemini 3.5 Flash Google0.6422026-05
3Meta logoMuse Spark (2026-04-08) Meta0.6212026-05
4OpenAI logoGPT-5.4 mini OpenAI0.5522026-05
5OpenAI logoGPT-5.4 (2026-03-05) OpenAI0.5382026-05
6Google logoGemini 2.5 Pro Google0.5292026-05
7DDeepSeek R1 DeepSeek0.4852026-05
8Anthropic logoClaude 4.6 Opus Anthropic0.4562026-05
9Anthropic logoClaude 3.7 Sonnet Anthropic0.452026-05
10Google logoGemini 2.0 Flash Google0.3422026-05

Scores appear exactly as MedHELM leaderboard (medhelm.org), v5.0.0 publishes them (official leaderboard). Version 5.0.0, last updated May 14, 2026, run by the Stanford-led maintainers on a roughly quarterly cadence. No Claude 5 family or GPT-5.6 rows yet. Mean win rate is relative to the evaluated cohort, so scores shift whenever the model set changes.

About the benchmark

publisherStanford CRFM / HAI and multi-institution collaborators
categorycomposite indices
released2025-02
size121 tasks / 31 datasets
scalemean win rate 0-1, higher better
result basisofficial leaderboard
sourceMedHELM leaderboard (medhelm.org), v5.0.0
last frontier result2026-05

What is MedHELM?

MedHELM is a composite benchmark from Stanford CRFM, released 2025-02: 121 tasks / 31 datasets, scored on a mean win rate 0-1 scale. Holistic evaluation of LLMs on 121 clinical tasks across 5 categories and 22 subcategories (31 datasets) in a clinician-validated taxonomy; ranked by mean win rate.

Which model leads MedHELM?

Gemini 3.1 Pro (Preview) (Google) holds the top current result on MedHELM at 0.652, per MedHELM leaderboard (medhelm.org), v5.0.0, as of 2026-05.

Where do the MedHELM numbers come from?

From MedHELM leaderboard (medhelm.org), v5.0.0 (official leaderboard). Version 5.0.0, last updated May 14, 2026, run by the Stanford-led maintainers on a roughly quarterly cadence. No Claude 5 family or GPT-5.6 rows yet. Mean win rate is relative to the evaluated cohort, so scores shift whenever the model set changes.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.