Clinical Benchmarks

HealthBench Professional: current results

OpenAI · 525 physician-authored tasks · index updated August 16, 2026

Claude Fable 5 holds the top current result on HealthBench Professional, 0.660 as of 2026-08, per healthbenchprofessional.com. 525 tasks that physicians picked out of 15,079 real workplace AI conversations, spanning care consults, clinical documentation, and medical research, each judged on a rubric physicians wrote for it.

This benchmark has a dedicated full leaderboard, with methodology and per-model pages, at healthbenchprofessional.com. The top of its table is mirrored below.

Current top results

Result detail

#modelscoreas of
1Anthropic logoClaude Fable 5 Anthropic0.6602026-08
2OpenAI logoGPT-5.6 Sol OpenAI0.6052026-08
3Anthropic logoClaude Opus 5 Anthropic0.5982026-08

Scores appear exactly as healthbenchprofessional.com publishes them (independently run). Grader GPT-5.4 at low reasoning effort with a length adjustment. Physician-written responses score 0.437 on the same rubrics.

About the benchmark

publisherOpenAI
categoryrubric-graded benchmarks
released2026-04
size525 physician-authored tasks
scale0 to 1, higher is better
result basisindependently run
sourcehealthbenchprofessional.com
last frontier result2026-08

What is HealthBench Professional?

HealthBench Professional is a rubric-graded benchmark from OpenAI, released 2026-04: 525 physician-authored tasks, scored on a 0 to 1 scale. 525 tasks that physicians picked out of 15,079 real workplace AI conversations, spanning care consults, clinical documentation, and medical research, each judged on a rubric physicians wrote for it.

Which model leads HealthBench Professional?

Claude Fable 5 (Anthropic) holds the top current result on HealthBench Professional at 0.660, per healthbenchprofessional.com, as of 2026-08.

Where do the HealthBench Professional numbers come from?

From healthbenchprofessional.com (independently run). Grader GPT-5.4 at low reasoning effort with a length adjustment. Physician-written responses score 0.437 on the same rubrics.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.