HealthBench Professional: current results
OpenAI · 525 physician-authored tasks · index updated August 16, 2026
Claude Fable 5 holds the top current result on HealthBench Professional, 0.660 as of 2026-08, per healthbenchprofessional.com. 525 tasks that physicians picked out of 15,079 real workplace AI conversations, spanning care consults, clinical documentation, and medical research, each judged on a rubric physicians wrote for it.
This benchmark has a dedicated full leaderboard, with methodology and per-model pages, at healthbenchprofessional.com. The top of its table is mirrored below.
Current top results
- 1
Claude Fable 50.660
- 2
GPT-5.6 Sol0.605
- 3
Claude Opus 50.598
Result detail
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Fable 5 Anthropic | 0.660 | 2026-08 | |
| 2 | GPT-5.6 Sol OpenAI | 0.605 | 2026-08 | |
| 3 | Claude Opus 5 Anthropic | 0.598 | 2026-08 | |
Scores appear exactly as healthbenchprofessional.com publishes them (independently run). Grader GPT-5.4 at low reasoning effort with a length adjustment. Physician-written responses score 0.437 on the same rubrics.
About the benchmark
| publisher | OpenAI |
|---|---|
| category | rubric-graded benchmarks |
| released | 2026-04 |
| size | 525 physician-authored tasks |
| scale | 0 to 1, higher is better |
| result basis | independently run |
| source | healthbenchprofessional.com |
| last frontier result | 2026-08 |
What is HealthBench Professional?
HealthBench Professional is a rubric-graded benchmark from OpenAI, released 2026-04: 525 physician-authored tasks, scored on a 0 to 1 scale. 525 tasks that physicians picked out of 15,079 real workplace AI conversations, spanning care consults, clinical documentation, and medical research, each judged on a rubric physicians wrote for it.
Which model leads HealthBench Professional?
Claude Fable 5 (Anthropic) holds the top current result on HealthBench Professional at 0.660, per healthbenchprofessional.com, as of 2026-08.
Where do the HealthBench Professional numbers come from?
From healthbenchprofessional.com (independently run). Grader GPT-5.4 at low reasoning effort with a length adjustment. Physician-written responses score 0.437 on the same rubrics.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.