Clinical Benchmarks

OpenAI logoOpenAI: healthcare benchmark results

19 models · 30 results · updated August 16, 2026

Everything the index holds for OpenAI, gathered in one place: GPT-5.6 Sol, GPT-5.5, GPT-5.4, GPT-5.6 Terra, GPT-5.6 Sol (August), Codex (GPT-5.6-sol), GPT-5.2-High, GPT-5.6 Luna, GPT-5.4 mini, GPT-5.4 (2026-03-05), Codex (GPT 5.5), GPT 5.1, GPT-5.4 (high reasoning), GPT-5.4 (low reasoning), GPT-4.1, GPT-4o, GPT-5.4 (computer-use agent), GPT-5.5 Instant (June Update), GPT-5.6 Luna (August). The lab holds first place on MAST (Medical AI Superintelligence Test), PhysicianBench, EHR-Complex, OpenAI Dynamic Mental Health Evaluations. Scores sit on each benchmark's own scale and never compare across rows from different boards.

Every result

Which healthcare benchmarks does OpenAI appear on?

As of August 16, 2026, OpenAI models hold 30 current results across 15 tracked benchmarks, through GPT-5.6 Sol, GPT-5.5, GPT-5.4, GPT-5.6 Terra, GPT-5.6 Sol (August), Codex (GPT-5.6-sol), GPT-5.2-High, GPT-5.6 Luna, GPT-5.4 mini, GPT-5.4 (2026-03-05), Codex (GPT 5.5), GPT 5.1, GPT-5.4 (high reasoning), GPT-5.4 (low reasoning), GPT-4.1, GPT-4o, GPT-5.4 (computer-use agent), GPT-5.5 Instant (June Update), GPT-5.6 Luna (August).

Where does OpenAI lead?

OpenAI models hold first place on MAST (Medical AI Superintelligence Test) (GPT-5.6 Sol, 60.2%); PhysicianBench (GPT-5.5, 46.3 ± 1.2); EHR-Complex (GPT-5.4 (high reasoning), 0.65); OpenAI Dynamic Mental Health Evaluations (GPT-5.5 Instant (June Update), 0.991).

Other labs with pages: Anthropic, Google, Moonshot AI, Alibaba, Meta, xAI. The full field is on the index.