Rubric-graded benchmarks
5 tracked · updated August 16, 2026
These benchmarks grade free-text answers against rubrics that physicians or domain experts wrote, criterion by criterion, instead of scoring a picked letter. They reward completeness, accuracy, and safe framing at once, which is why frontier scores on them run low and move slowly. All of them descend from the idea that a health answer is a judgment to be audited, not a fact to be matched.
HealthBench Professional
OpenAI · 525 tasks- 1
Claude Fable 50.660
- 2
GPT-5.6 Sol0.605
- 3
Claude Opus 50.598
HealthBench Hard
OpenAI · 1,000 conversations- 1
Muse Spark0.428
- 2
GPT-5.6 Sol0.331
- 3
GPT-5.6 Terra0.327
HealthBench
OpenAI · 5,000 conversations- 1
Claude Opus 567.1
- 2BBaichuan-M365.1
- 3
GPT-5.2-High63.3
- 4
GPT-5.6 Sol57.0
- 5
GPT-5.6 Terra57.0
Health Optimization Bench
healthoptimizationbench.com · 89 tasks- 1
Claude Fable 583.8
- 2
Grok 4.681.3
- 3
Claude Opus 578.3
WHBench
academic team · 47 scenarios- 1
Claude Opus 4.672.1%
- 2
Claude Sonnet 4.667.1%
- 3
GPT-5.466.8%
- 4
Gemini 3 Flash Preview64.7%
- 5
GPT-4.151.8%
Which rubric-graded benchmarks have current frontier-model results?
5 as of August 16, 2026: HealthBench Professional (Claude Fable 5 leads at 0.660); HealthBench Hard (Muse Spark leads at 0.428); HealthBench (Claude Opus 5 leads at 67.1); Health Optimization Bench (Claude Fable 5 leads at 83.8); WHBench (Claude Opus 4.6 leads at 72.1%).
The other categories sit on the index: agentic and workflow benchmarks, documentation and coding benchmarks, safety benchmarks, knowledge and exam benchmarks, composite indices.