OpenAI
OpenAI models with results here, each placed on its own board.
- Models
- 25
- Results
- 83
- Boards
- 13 of 13
Models
| Model | Kind | Released | Boards |
|---|---|---|---|
| GPT-6.1 Sol | Model | 29 Sep 2026 | 2 |
| GPT-6 Luna | Model | 22 Sep 2026 | 3 |
| GPT-6 Sol | Model | 22 Sep 2026 | 3 |
| GPT-6 Astra | Model | 3 Sep 2026 | 3 |
| GPT-5.6 Luna | Model | 9 Jul 2026 | 4 |
| GPT-5.6 Sol | Model | 9 Jul 2026 | 8 |
| GPT-5.6 Terra | Model | 9 Jul 2026 | 4 |
| GPT-5.5 Instant | Model | 5 May 2026 | 1 |
| GPT-5.5 | Model | 23 Apr 2026 | 7 |
| GPT-5.4 mini | Model | 17 Mar 2026 | 3 |
| GPT-5.4 nano | Model | 17 Mar 2026 | 2 |
| GPT-5.4 | Model | 5 Mar 2026 | 10 |
| GPT-5.3-Codex | Model | 5 Feb 2026 | 1 |
| GPT-5.2 | Model | 11 Dec 2025 | 5 |
| GPT-5.1 | Model | 12 Nov 2025 | 3 |
| GPT-5 | Model | 7 Aug 2025 | 5 |
| GPT-5 mini | Model | 7 Aug 2025 | 2 |
| GPT-5 nano | Model | 7 Aug 2025 | 2 |
| GPT OSS 120B | Model | 5 Aug 2025 | 1 |
| GPT OSS 20B | Model | 5 Aug 2025 | 1 |
| o3 | Model | 16 Apr 2025 | 2 |
| o4-mini | Model | 16 Apr 2025 | 2 |
| GPT-4.1 | Model | 14 Apr 2025 | 1 |
| GPT-4.1 mini | Model | 14 Apr 2025 | 1 |
| GPT-4o | Model | 13 May 2024 | 1 |
Results
One line per result. Dark tick: the board leader.
Clinical reasoning and knowledge
- HealthBench ProfessionalAnthropic public-API reproduction; max effort; no system prompt; Opus 4.8 grader; length-adjusted; raw 74.0%0.703Rank 1 of 31Leads this board
- HealthBench Professionallength-adjusted, max reasoning effort (69.5 unadjusted, 4,097 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'.0.647Rank 5 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench ProfessionalOpenAI; length-adjusted; maximum reasoning effort; 60.8 (61.2, 2119) (adjusted, raw, mean response characters)0.608Rank 8 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench ProfessionalOpenAI; length-adjusted; maximum reasoning effort; 60.8 (59.5, 1573) (adjusted, raw, mean response characters)0.608Rank 8 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492.0.605Rank 10 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.0.577Rank 14 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493.0.557Rank 18 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench ProfessionalChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars)0.540Rank 20 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (57.2 unadjusted, 3818 chars)0.518Rank 22 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars)0.481Rank 24 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars)0.462Rank 25 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars)0.459Rank 26 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench ProfessionalChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars)0.441Rank 28 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars)0.396Rank 29 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- HealthBench Professionallength-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.0.384Rank 30 of 31Leader GPT-6 Astra (Anthropic run) 0.703
- 81.5Rank 1 of 22 here, 5 models on the boardLeads this board
- MedXpertQA (MM)Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)77.1Rank 6 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 73.3Rank 8 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
- 60.2Rank 1 of 8 here, 11 models on the boardLeads this board
- 0.552Rank 4 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
- 0.538Rank 5 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
Documentation and coding
- 88.09Rank 12 of 105Leader Claude Opus 5.5 91.43
- 87.91Rank 14 of 105Leader Claude Opus 5.5 91.43
- 86.87Rank 17 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID openai/gpt-6.1-sol; max_output_tokens=128000; reasoning_effort=max86.45Rank 20 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID openai/gpt-5.6-sol; max_output_tokens=30000; reasoning_effort=max85.23Rank 28 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID openai/gpt-5.6-luna; max_output_tokens=30000; reasoning_effort=max84.39Rank 33 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID openai/gpt-5.2-2025-12-11; max_output_tokens=30000; reasoning_effort=xhigh84.39Rank 33 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID openai/gpt-6-luna; max_output_tokens=128000; reasoning_effort=max83.71Rank 40 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID openai/gpt-5-2025-08-07; max_output_tokens=30000; reasoning_effort=high83.65Rank 41 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID openai/gpt-5.6-terra; max_output_tokens=30000; reasoning_effort=xhigh82.87Rank 47 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID openai/gpt-6-sol; max_output_tokens=128000; reasoning_effort=max82.03Rank 49 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID openai/gpt-5-mini-2025-08-07; max_output_tokens=30000; reasoning_effort=high80.58Rank 53 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID openai/gpt-5.4-2026-03-05; max_output_tokens=30000; reasoning_effort=xhigh77.55Rank 65 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID openai/gpt-5.4-nano-2026-03-17; max_output_tokens=30000; reasoning_effort=high77.09Rank 68 of 105Leader Claude Opus 5.5 91.43
- 76.65Rank 70 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID openai/gpt-5-nano-2025-08-07; max_output_tokens=30000; reasoning_effort=high72.86Rank 81 of 105Leader Claude Opus 5.5 91.43
- MedScribe (Vals AI)model ID openai/o4-mini-2025-04-16; max_output_tokens=30000; reasoning_effort=high69.14Rank 93 of 105Leader Claude Opus 5.5 91.43
- 52.73Rank 12 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID openai/gpt-5.2-2025-12-11; reasoning_effort=xhigh; max_output_tokens=3000049.75Rank 17 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID openai/gpt-5-2025-08-07; reasoning_effort=high; max_output_tokens=3000049.63Rank 18 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID openai/gpt-5.5; reasoning_effort=xhigh; temperature=1; max_output_tokens=3000049.10Rank 23 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID openai/gpt-6.1-sol; reasoning_effort=max; max_output_tokens=12800048.84Rank 25 of 103Leader Claude Opus 5 63.57
- 48.49Rank 26 of 103Leader Claude Opus 5 63.57
- 47.29Rank 31 of 103Leader Claude Opus 5 63.57
- 47.07Rank 33 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=12800044.69Rank 38 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID openai/gpt-5.6-sol; reasoning_effort=max; max_output_tokens=3000043.97Rank 40 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID openai/gpt-5.6-terra; reasoning_effort=xhigh; max_output_tokens=3000043.41Rank 42 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID openai/gpt-5-mini-2025-08-07; reasoning_effort=high; max_output_tokens=3000043.05Rank 45 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID openai/gpt-5.6-luna; reasoning_effort=max; max_output_tokens=3000042.39Rank 48 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID openai/gpt-5.4-2026-03-05; reasoning_effort=xhigh; max_output_tokens=3000041.29Rank 52 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID openai/gpt-5.4-nano-2026-03-17; reasoning_effort=high; max_output_tokens=3000041.03Rank 56 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID openai/o4-mini-2025-04-16; reasoning_effort=high; max_output_tokens=3000033.79Rank 80 of 103Leader Claude Opus 5 63.57
- MedCode (Vals AI)model ID openai/gpt-5-nano-2025-08-07; reasoning_effort=high; max_output_tokens=3000030.44Rank 92 of 103Leader Claude Opus 5 63.57
EHR and workflow agents
- 46.3Rank 7 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- 27.7Rank 12 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
- EHR-Complexaverage over 12 intent columns; run as human-validation configuration, not in the headline 12-model table0.650Rank 1 of 18Leads this board
- 0.580Rank 6 of 18Leader GPT-5.4 (high reasoning) 0.650
- EHR-ComplexTable 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns0.490Rank 10 of 18Leader GPT-5.4 (high reasoning) 0.650
- EHR-ComplexTable 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns0.470Rank 11 of 18Leader GPT-5.4 (high reasoning) 0.650
- EHR-ComplexTable 3; macro-average across 12 intent/scope columns; temperature 0; up to 50 SQL/Python interaction turns0.310Rank 15 of 18Leader GPT-5.4 (high reasoning) 0.650
- 45Rank 2 of 12Leader Claude Code (Opus 5) 55
- 42Rank 3 of 12Leader Claude Code (Opus 5) 55
- 35Rank 5 of 12Leader Claude Code (Opus 5) 55
- 28Rank 7 of 12Leader Claude Code (Opus 5) 55
- 22Rank 9 of 12Leader Claude Code (Opus 5) 55
- 16Rank 12 of 12Leader Claude Code (Opus 5) 55
- 25.3Rank 7 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 20.9Rank 12 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 16.0Rank 20 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 13.3Rank 26 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 13.3Rank 26 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 8.4Rank 34 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
- 26.7Rank 2 of 7 here, 5 models on the boardLeader Claude Opus 4.6 (computer-use agent) 36.3
- HealthAdminBenchscreenshot-only, detailed prompting; authors' standardized harness, no native CUA5.9Rank 7 of 7 here, 5 models on the boardLeader Claude Opus 4.6 (computer-use agent) 36.3
Safety
- 75.6Rank 2 of 28Leader Gemini-3.1-Pro 80.7
- 70.7Rank 5 of 28Leader Gemini-3.1-Pro 80.7
- 68.1Rank 8 of 28Leader Gemini-3.1-Pro 80.7
- 59.1Rank 14 of 28Leader Gemini-3.1-Pro 80.7
- 70.1Rank 8 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
- 70.0Rank 9 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
- First, Do NOHARM (v2)from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI); not in the Latest Flagships ranking68.6Rank 10 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2