Clinical Benchmarks

Anthropic logoClaude Opus 4.6: healthcare benchmark results

Anthropic · 5 boards · updated August 16, 2026

The index currently holds 5 results for Claude Opus 4.6: 0.456 on MedHELM (8 of 11), 86.74% on MedScribe (Vals AI) (6 of 84), 64.8% on MedXpertQA (MM) (7 of 8), 31.7 ± 2.3 on PhysicianBench (2 of 13), 72.1% on WHBench (1 of 22). It tops WHBench. Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscorepositionas of
MedHELM
via MedHELM leaderboard (medhelm.org), v5.0.0
0.4568 of 112026-05
MedScribe (Vals AI)
via Vals AI MedScribe leaderboard
86.74%6 of 842026-02
MedXpertQA (MM)
via benchlm.ai mirror of Meta's Muse Spark evaluation
64.8%7 of 82026-08
PhysicianBench
via PhysicianBench paper (Table 2)
31.7 ± 2.32 of 132026-05
WHBench
via arXiv paper (v2 revised 2026-07-23)
72.1%1 of 222026-07

Position counts against the source's full board, including rows this index does not mirror. Config caveats, where a source noted any, are on each benchmark's page. Sources also list this model as "Claude 4.6 Opus".

Which healthcare benchmarks is Claude Opus 4.6 scored on?

As of August 16, 2026, Claude Opus 4.6 holds current results on 5 tracked benchmarks: MedHELM, MedScribe (Vals AI), MedXpertQA (MM), PhysicianBench, WHBench.

How does Claude Opus 4.6 rank on them?

Claude Opus 4.6 stands at 0.456 on MedHELM (8 of 11), 86.74% on MedScribe (Vals AI) (6 of 84), 64.8% on MedXpertQA (MM) (7 of 8), 31.7 ± 2.3 on PhysicianBench (2 of 13), 72.1% on WHBench (1 of 22). It holds first place on WHBench.

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.