Clinical Benchmarks

OpenAI

OpenAI models with results here, each placed on its own board.

Models
25
Results
83
Boards
13 of 13

openai.com

Models

ModelKindReleasedBoards
GPT-6.1 SolModel29 Sep 20262
GPT-6 LunaModel22 Sep 20263
GPT-6 SolModel22 Sep 20263
GPT-6 AstraModel3 Sep 20263
GPT-5.6 LunaModel9 Jul 20264
GPT-5.6 SolModel9 Jul 20268
GPT-5.6 TerraModel9 Jul 20264
GPT-5.5 InstantModel5 May 20261
GPT-5.5Model23 Apr 20267
GPT-5.4 miniModel17 Mar 20263
GPT-5.4 nanoModel17 Mar 20262
GPT-5.4Model5 Mar 202610
GPT-5.3-CodexModel5 Feb 20261
GPT-5.2Model11 Dec 20255
GPT-5.1Model12 Nov 20253
GPT-5Model7 Aug 20255
GPT-5 miniModel7 Aug 20252
GPT-5 nanoModel7 Aug 20252
GPT OSS 120BModel5 Aug 20251
GPT OSS 20BModel5 Aug 20251
o3Model16 Apr 20252
o4-miniModel16 Apr 20252
GPT-4.1Model14 Apr 20251
GPT-4.1 miniModel14 Apr 20251
GPT-4oModel13 May 20241

Results

One line per result. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. HealthBench Professional
    Anthropic public-API reproduction; max effort; no system prompt; Opus 4.8 grader; length-adjusted; raw 74.0%
    0.703
    Rank 1 of 31Leads this board
    Independent run
    Measured Sep 2026
  2. HealthBench Professional
    length-adjusted, max reasoning effort (69.5 unadjusted, 4,097 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'.
    0.647
    Rank 5 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Sep 2026
  3. HealthBench Professional
    OpenAI; length-adjusted; maximum reasoning effort; 60.8 (61.2, 2119) (adjusted, raw, mean response characters)
    0.608
    Rank 8 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Sep 2026
  4. HealthBench Professional
    OpenAI; length-adjusted; maximum reasoning effort; 60.8 (59.5, 1573) (adjusted, raw, mean response characters)
    0.608
    Rank 8 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Sep 2026
  5. HealthBench Professional
    length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492.
    0.605
    Rank 10 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026
  6. HealthBench Professional
    length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
    0.577
    Rank 14 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026
  7. HealthBench Professional
    length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493.
    0.557
    Rank 18 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026
  8. HealthBench Professional
    ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars)
    0.540
    Rank 20 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Aug 2026
  9. HealthBench Professional
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (57.2 unadjusted, 3818 chars)
    0.518
    Rank 22 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026
  10. HealthBench Professional
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars)
    0.481
    Rank 24 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026
  11. HealthBench Professional
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars)
    0.462
    Rank 25 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026
  12. HealthBench Professional
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars)
    0.459
    Rank 26 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026
  13. HealthBench Professional
    ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars)
    0.441
    Rank 28 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Aug 2026
  14. HealthBench Professional
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars)
    0.396
    Rank 29 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026
  15. HealthBench Professional
    length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
    0.384
    Rank 30 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured May 2026
  16. MedXpertQA (MM)
    Qwen-run comparison in the Qwen3.8-Max launch post
    81.5
    Rank 1 of 22 here, 5 models on the boardLeads this board
    Independent run
  17. MedXpertQA (MM)
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    77.1
    Rank 6 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Independent run
    Measured Apr 2026
  18. MedXpertQA (MM)
    Qwen-run comparison in the Qwen3.5-397B-A17B model card
    73.3
    Rank 8 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Independent run
  19. MAST (Medical AI Superintelligence Test)
    MAST in preview; 'exact scores may change'
    60.2
    Rank 1 of 8 here, 11 models on the boardLeads this board
    Official leaderboard
    Measured Aug 2026
  20. 0.552
    Rank 4 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
    Official leaderboard
    Measured May 2026
  21. 0.538
    Rank 5 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
    Official leaderboard
    Measured May 2026

Documentation and coding

  1. 88.09
    Rank 12 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. 87.91
    Rank 14 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  3. 86.87
    Rank 17 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  4. MedScribe (Vals AI)
    model ID openai/gpt-6.1-sol; max_output_tokens=128000; reasoning_effort=max
    86.45
    Rank 20 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  5. MedScribe (Vals AI)
    model ID openai/gpt-5.6-sol; max_output_tokens=30000; reasoning_effort=max
    85.23
    Rank 28 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  6. MedScribe (Vals AI)
    model ID openai/gpt-5.6-luna; max_output_tokens=30000; reasoning_effort=max
    84.39
    Rank 33 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  7. MedScribe (Vals AI)
    model ID openai/gpt-5.2-2025-12-11; max_output_tokens=30000; reasoning_effort=xhigh
    84.39
    Rank 33 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  8. MedScribe (Vals AI)
    model ID openai/gpt-6-luna; max_output_tokens=128000; reasoning_effort=max
    83.71
    Rank 40 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  9. MedScribe (Vals AI)
    model ID openai/gpt-5-2025-08-07; max_output_tokens=30000; reasoning_effort=high
    83.65
    Rank 41 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  10. MedScribe (Vals AI)
    model ID openai/gpt-5.6-terra; max_output_tokens=30000; reasoning_effort=xhigh
    82.87
    Rank 47 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  11. MedScribe (Vals AI)
    model ID openai/gpt-6-sol; max_output_tokens=128000; reasoning_effort=max
    82.03
    Rank 49 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  12. MedScribe (Vals AI)
    model ID openai/gpt-5-mini-2025-08-07; max_output_tokens=30000; reasoning_effort=high
    80.58
    Rank 53 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  13. MedScribe (Vals AI)
    model ID openai/gpt-5.4-2026-03-05; max_output_tokens=30000; reasoning_effort=xhigh
    77.55
    Rank 65 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  14. MedScribe (Vals AI)
    model ID openai/gpt-5.4-nano-2026-03-17; max_output_tokens=30000; reasoning_effort=high
    77.09
    Rank 68 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  15. MedScribe (Vals AI)
    model ID openai/o3-2025-04-16; max_output_tokens=30000; reasoning_effort=high
    76.65
    Rank 70 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  16. MedScribe (Vals AI)
    model ID openai/gpt-5-nano-2025-08-07; max_output_tokens=30000; reasoning_effort=high
    72.86
    Rank 81 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  17. MedScribe (Vals AI)
    model ID openai/o4-mini-2025-04-16; max_output_tokens=30000; reasoning_effort=high
    69.14
    Rank 93 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  18. 52.73
    Rank 12 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  19. MedCode (Vals AI)
    model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=30000
    49.75
    Rank 17 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  20. MedCode (Vals AI)
    model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=30000
    49.63
    Rank 18 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  21. MedCode (Vals AI)
    model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=30000
    49.10
    Rank 23 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  22. MedCode (Vals AI)
    model ID openai/gpt-6.1-sol; reasoning_effort=max; max_output_tokens=128000
    48.84
    Rank 25 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  23. 48.49
    Rank 26 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  24. MedCode (Vals AI)
    model ID openai/o3-2025-04-16; reasoning_effort=high; max_output_tokens=30000
    47.29
    Rank 31 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  25. MedCode (Vals AI)
    model ID openai/gpt-6-sol; reasoning_effort=max; max_output_tokens=128000
    47.07
    Rank 33 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  26. MedCode (Vals AI)
    model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000
    44.69
    Rank 38 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  27. MedCode (Vals AI)
    model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=30000
    43.97
    Rank 40 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  28. MedCode (Vals AI)
    model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=30000
    43.41
    Rank 42 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  29. MedCode (Vals AI)
    model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=30000
    43.05
    Rank 45 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  30. MedCode (Vals AI)
    model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=30000
    42.39
    Rank 48 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  31. MedCode (Vals AI)
    model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=30000
    41.29
    Rank 52 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  32. MedCode (Vals AI)
    model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=30000
    41.03
    Rank 56 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  33. MedCode (Vals AI)
    model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=30000
    33.79
    Rank 80 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  34. MedCode (Vals AI)
    model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=30000
    30.44
    Rank 92 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. PhysicianBench
    pass@1; Pass^3 28.0
    46.3
    Rank 7 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Official leaderboard
    Measured May 2026
  2. 27.7
    Rank 12 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Official leaderboard
    Measured May 2026
  3. EHR-Complex
    average over 12 intent columns; run as human-validation configuration, not in the headline 12-model table
    0.650
    Rank 1 of 18Leads this board
    Official leaderboard
    Measured Jun 2026
  4. EHR-Complex
    validation configuration
    0.580
    Rank 6 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  5. EHR-Complex
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    0.490
    Rank 10 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  6. EHR-Complex
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    0.470
    Rank 11 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  7. EHR-Complex
    Table 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns
    0.310
    Rank 15 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  8. 45
    Rank 2 of 12Leader Claude Code (Opus 5) 55
    Official leaderboard
    Measured Jul 2026
  9. 42
    Rank 3 of 12Leader Claude Code (Opus 5) 55
    Official leaderboard
    Measured Jul 2026
  10. 35
    Rank 5 of 12Leader Claude Code (Opus 5) 55
    Official leaderboard
    Measured Jul 2026
  11. 28
    Rank 7 of 12Leader Claude Code (Opus 5) 55
    Official leaderboard
    Measured Jul 2026
  12. HealthAgentBench
    $1.0/task; Codex harness; mean success across 3 attempts × 54 tasks
    22
    Rank 9 of 12Leader Claude Code (Opus 5) 55
    Official leaderboard
    Measured Jul 2026
  13. 16
    Rank 12 of 12Leader Claude Code (Opus 5) 55
    Official leaderboard
    Measured Jul 2026
  14. 25.3
    Rank 7 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Jul 2026
  15. 20.9
    Rank 12 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Apr 2026
  16. CHI-Bench
    All Domains pass@1; codex harness
    16.0
    Rank 20 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Apr 2026
  17. CHI-Bench
    All Domains pass@1; codex harness
    13.3
    Rank 26 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Jul 2026
  18. CHI-Bench
    All Domains pass@1; codex harness
    13.3
    Rank 26 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Jul 2026
  19. CHI-Bench
    All Domains pass@1; codex harness
    8.4
    Rank 34 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  20. HealthAdminBench
    screenshot-only, detailed prompting; subtask rate 82.8%
    26.7
    Rank 2 of 7 here, 5 models on the boardLeader Claude Opus 4.6 (computer-use agent) 36.3
    Official leaderboard
    Measured Apr 2026
  21. HealthAdminBench
    screenshot-only, detailed prompting; authors' standardized harness, no native CUA
    5.9
    Rank 7 of 7 here, 5 models on the boardLeader Claude Opus 4.6 (computer-use agent) 36.3
    Official leaderboard
    Measured Apr 2026

Safety

  1. MedPIC
    zero-shot; independent questions; exact option-set match
    75.6
    Rank 2 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  2. MedPIC
    zero-shot; independent questions; exact option-set match
    70.7
    Rank 5 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  3. MedPIC
    zero-shot; independent questions; exact option-set match
    68.1
    Rank 8 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  4. MedPIC
    zero-shot; independent questions; exact option-set match
    59.1
    Rank 14 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  5. 70.1
    Rank 8 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026
  6. 70.0
    Rank 9 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026
  7. First, Do NOHARM (v2)
    from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI); not in the Latest Flagships ranking
    68.6
    Rank 10 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026