Clinical Benchmarks

OpenAI logoGPT-5.6 Sol: healthcare benchmark results

OpenAI · 5 boards · updated August 16, 2026

The index currently holds 5 results for GPT-5.6 Sol: 0.605 on HealthBench Professional (2 of 9), 0.331 on HealthBench Hard (2 of 9), 57.0 on HealthBench (4 of 8), 60.2% on MAST (Medical AI Superintelligence Test) (1 of 11), 70.1% on First, Do NOHARM (v2) (3 of 19). It tops MAST (Medical AI Superintelligence Test). Scores below sit on different scales and come from different graders, so read each against its own benchmark, never against the others.

Results by benchmark

benchmarkscorepositionas of
HealthBench Professional
via healthbenchprofessional.com
0.6052 of 92026-08
HealthBench Hard
via healthbenchhard.ai
0.3312 of 92026-08
HealthBench
via OpenAI Deployment Safety Hub (GPT-5.6 system card + August 2026 updates); benchlm.ai and llm-stats.com mirror
57.04 of 82026-06
MAST (Medical AI Superintelligence Test)
via ARISE MAST leaderboard
60.2%1 of 112026-08
First, Do NOHARM (v2)
via ARISE MAST technical leaderboard
70.1%3 of 192026-08

Position counts against the source's full board, including rows this index does not mirror. Config caveats, where a source noted any, are on each benchmark's page.

Which healthcare benchmarks is GPT-5.6 Sol scored on?

As of August 16, 2026, GPT-5.6 Sol holds current results on 5 tracked benchmarks: HealthBench Professional, HealthBench Hard, HealthBench, MAST (Medical AI Superintelligence Test), First, Do NOHARM (v2).

How does GPT-5.6 Sol rank on them?

GPT-5.6 Sol stands at 0.605 on HealthBench Professional (2 of 9), 0.331 on HealthBench Hard (2 of 9), 57.0 on HealthBench (4 of 8), 60.2% on MAST (Medical AI Superintelligence Test) (1 of 11), 70.1% on First, Do NOHARM (v2) (3 of 19). It holds first place on MAST (Medical AI Superintelligence Test).

The benchmarks themselves are described on their pages, linked in the table above, and the whole field is on the index.