GPT-5.2: healthcare benchmark results
OpenAI · 5 boards · updated September 8, 2026
The index currently holds 5 results for GPT-5.2: 0.459 on HealthBench Professional (17 of 22 indexed rows), 0.343 on HealthBench Hard (4 of 16 indexed rows), 63.3 on HealthBench (3 of 18 indexed rows), 73.3 on MedXpertQA (MM) (7 of 15 indexed rows), 0.975 on OpenAI Dynamic Mental Health Evaluations (14 of 15 indexed rows). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.
Results by benchmark
| benchmark | score | index position | as of |
|---|---|---|---|
| HealthBench Professional | 0.459 | 17 of 22 | 2026-06 |
| HealthBench Hard | 0.343 | 4 of 16 | 2026-06 |
| HealthBench via paper, arxiv.org | 63.3 | 3 of 18 | 2026-02 |
| MedXpertQA (MM) | 73.3 | 7 of 15 | |
| OpenAI Dynamic Mental Health Evaluations | 0.975 | 14 of 15 | 2026-03 |
Position follows the rows held in this index, including separate configurations. A source board may contain additional rows; its size is listed separately when it differs. Config caveats, where a source noted any, are on each benchmark's page, and the document behind each score is cited on the sources page. Sources also list this model as "GPT-5.2-High" and "GPT-5.2 Thinking".
Which healthcare benchmarks is GPT-5.2 scored on?
As of September 8, 2026, GPT-5.2 holds current results on 5 tracked benchmarks: HealthBench Professional, HealthBench Hard, HealthBench, MedXpertQA (MM), OpenAI Dynamic Mental Health Evaluations.
How does GPT-5.2 rank on them?
GPT-5.2 stands at 0.459 on HealthBench Professional (17 of 22 indexed rows), 0.343 on HealthBench Hard (4 of 16 indexed rows), 63.3 on HealthBench (3 of 18 indexed rows), 73.3 on MedXpertQA (MM) (7 of 15 indexed rows), 0.975 on OpenAI Dynamic Mental Health Evaluations (14 of 15 indexed rows).
The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.