WHBench: current results
Independent researchers (Maurya, Govindgari, Kumar) · 47 scenarios / 3,100 scored responses across 22 models · index updated August 16, 2026
Claude Opus 4.6 holds the top current result on WHBench, 72.1% as of 2026-07, per arXiv paper (v2 revised 2026-07-23). Women's health: 47 expert-crafted scenarios across 10 topics graded on a 23-criterion rubric for clinical accuracy, safety, equity, and guideline adherence; targets failure modes like outdated guidelines, unsafe omissions, dosing errors, equity blind spots.
Current results
- 1
Claude Opus 4.672.1%
- 2
Claude Sonnet 4.667.1%
- 3
GPT-5.466.8%
- 4
Gemini 3 Flash Preview64.7%
- 5
GPT-4.151.8%
- 6
GPT-4o44.6%
Result detail
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | Claude Opus 4.6 Anthropic 95% CI 69.6-74.4; evaluations run March 2026 | 72.1% | 2026-07 | |
| 2 | Claude Sonnet 4.6 Anthropic 95% CI 64.5-69.6 | 67.1% | 2026-07 | |
| 3 | GPT-5.4 OpenAI 95% CI 64.5-69.2 | 66.8% | 2026-07 | |
| 4 | Gemini 3 Flash Preview Google | 64.7% | 2026-07 | |
| 5 | GPT-4.1 OpenAI | 51.8% | 2026-07 | |
| 6 | GPT-4o OpenAI | 44.6% | 2026-07 | |
Scores appear exactly as arXiv paper (v2 revised 2026-07-23) publishes them (independently run). An academic study with expert validation rather than a live leaderboard; the model set was frozen in March 2026, before GPT-5.6 and the Claude 5 family shipped.
About the benchmark
| publisher | Independent researchers (Maurya, Govindgari, Kumar) |
|---|---|
| category | rubric-graded benchmarks |
| released | 2026-04 |
| size | 47 scenarios / 3,100 scored responses across 22 models |
| scale | mean normalized percentage 0-100, higher better |
| result basis | independently run |
| source | arXiv paper (v2 revised 2026-07-23) |
| last frontier result | 2026-07 |
What is WHBench?
WHBench is a rubric-graded benchmark from academic team, released 2026-04: 47 scenarios / 3,100 scored responses across 22 models, scored on a mean normalized percentage 0-100 scale. Women's health: 47 expert-crafted scenarios across 10 topics graded on a 23-criterion rubric for clinical accuracy, safety, equity, and guideline adherence; targets failure modes like outdated guidelines, unsafe omissions, dosing errors, equity blind spots.
Which model leads WHBench?
Claude Opus 4.6 (Anthropic) holds the top current result on WHBench at 72.1%, per arXiv paper (v2 revised 2026-07-23), as of 2026-07.
Where do the WHBench numbers come from?
From arXiv paper (v2 revised 2026-07-23) (independently run). An academic study with expert validation rather than a live leaderboard; the model set was frozen in March 2026, before GPT-5.6 and the Claude 5 family shipped.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.