Clinical Benchmarks

HealthBench Hard: current results

OpenAI · 1,000 conversations · index updated August 16, 2026

Muse Spark holds the top current result on HealthBench Hard, 0.428 as of 2026-08, per healthbenchhard.ai. The bottom fifth of HealthBench: 1,000 conversations where frontier models failed most at the May 2025 release, still graded on the original physician-written rubrics.

This benchmark has a dedicated full leaderboard, with methodology and per-model pages, at healthbenchhard.ai. The top of its table is mirrored below.

Current top results

Result detail

#modelscoreas of
1Meta logoMuse Spark Meta0.4282026-08
2OpenAI logoGPT-5.6 Sol OpenAI0.3312026-08
3OpenAI logoGPT-5.6 Terra OpenAI0.3272026-08

Scores appear exactly as healthbenchhard.ai publishes them (mixed sources). Default grader GPT-4.1. Best score at the May 2025 release was o3's 0.320.

About the benchmark

publisherOpenAI
categoryrubric-graded benchmarks
released2025-05
size1,000 conversations
scale0 to 1, higher is better
result basismixed sources
sourcehealthbenchhard.ai
last frontier result2026-08

What is HealthBench Hard?

HealthBench Hard is a rubric-graded benchmark from OpenAI, released 2025-05: 1,000 conversations, scored on a 0 to 1 scale. The bottom fifth of HealthBench: 1,000 conversations where frontier models failed most at the May 2025 release, still graded on the original physician-written rubrics.

Which model leads HealthBench Hard?

Muse Spark (Meta) holds the top current result on HealthBench Hard at 0.428, per healthbenchhard.ai, as of 2026-08.

Where do the HealthBench Hard numbers come from?

From healthbenchhard.ai (mixed sources). Default grader GPT-4.1. Best score at the May 2025 release was o3's 0.320.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.