OpenAI: healthcare benchmark results
19 models · 30 results · updated August 16, 2026
Everything the index holds for OpenAI, gathered in one place: GPT-5.6 Sol, GPT-5.5, GPT-5.4, GPT-5.6 Terra, GPT-5.6 Sol (August), Codex (GPT-5.6-sol), GPT-5.2-High, GPT-5.6 Luna, GPT-5.4 mini, GPT-5.4 (2026-03-05), Codex (GPT 5.5), GPT 5.1, GPT-5.4 (high reasoning), GPT-5.4 (low reasoning), GPT-4.1, GPT-4o, GPT-5.4 (computer-use agent), GPT-5.5 Instant (June Update), GPT-5.6 Luna (August). The lab holds first place on MAST (Medical AI Superintelligence Test), PhysicianBench, EHR-Complex, OpenAI Dynamic Mental Health Evaluations. Scores sit on each benchmark's own scale and never compare across rows from different boards.
Every result
Which healthcare benchmarks does OpenAI appear on?
As of August 16, 2026, OpenAI models hold 30 current results across 15 tracked benchmarks, through GPT-5.6 Sol, GPT-5.5, GPT-5.4, GPT-5.6 Terra, GPT-5.6 Sol (August), Codex (GPT-5.6-sol), GPT-5.2-High, GPT-5.6 Luna, GPT-5.4 mini, GPT-5.4 (2026-03-05), Codex (GPT 5.5), GPT 5.1, GPT-5.4 (high reasoning), GPT-5.4 (low reasoning), GPT-4.1, GPT-4o, GPT-5.4 (computer-use agent), GPT-5.5 Instant (June Update), GPT-5.6 Luna (August).
Where does OpenAI lead?
OpenAI models hold first place on MAST (Medical AI Superintelligence Test) (GPT-5.6 Sol, 60.2%); PhysicianBench (GPT-5.5, 46.3 ± 1.2); EHR-Complex (GPT-5.4 (high reasoning), 0.65); OpenAI Dynamic Mental Health Evaluations (GPT-5.5 Instant (June Update), 0.991).
Other labs with pages: Anthropic, Google, Moonshot AI, Alibaba, Meta, xAI. The full field is on the index.