GPT-5.4: healthcare benchmark results
OpenAI · 3 boards · updated August 16, 2026
The index currently holds 3 results for GPT-5.4: 77.1% on MedXpertQA (MM) (4 of 8), 27.7 ± 1.5 on PhysicianBench (4 of 13), 66.8% on WHBench (3 of 22). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | position | as of |
|---|---|---|---|
| MedXpertQA (MM) via benchlm.ai mirror of Meta's Muse Spark evaluation | 77.1% | 4 of 8 | 2026-08 |
| PhysicianBench via PhysicianBench paper (Table 2) | 27.7 ± 1.5 | 4 of 13 | 2026-05 |
| WHBench via arXiv paper (v2 revised 2026-07-23) | 66.8% | 3 of 22 | 2026-07 |
Position counts against the source's full board, including rows this index does not mirror. Config caveats, where a source noted any, are on each benchmark's page.
Which healthcare benchmarks is GPT-5.4 scored on?
As of August 16, 2026, GPT-5.4 holds current results on 3 tracked benchmarks: MedXpertQA (MM), PhysicianBench, WHBench.
How does GPT-5.4 rank on them?
GPT-5.4 stands at 77.1% on MedXpertQA (MM) (4 of 8), 27.7 ± 1.5 on PhysicianBench (4 of 13), 66.8% on WHBench (3 of 22).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.