Clinical Benchmarks

GPT-6 Luna

Released 22 Sep 2026$0.1 input, $0.5 output per million tokens1.05M contextproprietaryAlso written as openai/gpt-6-lunaCompare with other models

Clinical Benchmarks Index
73.7rank 19 of 148; 73.7 × 1 = 73.7, from 3 of 10 boards
Boards
3 of 13
Results
3
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. HealthBench Professional
    OpenAI; length-adjusted; maximum reasoning effort; 60.8 (61.2, 2119) (adjusted, raw, mean response characters)
    0.608
    Rank 8 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Sep 2026

Documentation and coding

  1. MedScribe (Vals AI)
    model ID openai/gpt-6-luna; max_output_tokens=128000; reasoning_effort=max
    83.71
    Rank 40 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000
    44.69
    Rank 38 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 8HealthBench Professional OpenAI; length-adjusted; maximum reasoning effort; 60.8 (61.2, 2119) (adjusted,… 0.608
    Printed as 60.8Vendor-reported, measured Sep 2026Configuration: OpenAI; length-adjusted; maximum reasoning effort; 60.8 (61.2, 2119) (adjusted, raw, mean response characters)
    GPT-6 Astra System Card - HealthBench (Deployment Safety Hub) system card, OpenAI, 3 Sep 2026. Section 11.4.1, Table 29, HealthBench Professional length-adjusted, GPT-6 Luna column; September 22 correction.
    Evaluation | GPT-5.5 | GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna HealthBench Professional length-adjusted | 51.8 (57.2, 3818) | 60.5 (64.1, 3228) | 57.7 (62.4, 3618) | 55.7 (59.8, 3389) | 64.7 (68.2, 3185) | 60.8 (59.5, 1573) | 60.8 (61.2, 2119)
    Every result from this document
  2. 40MedScribe (Vals AI) model ID openai/gpt-6-luna; max_output_tokens=128000; reasoning_effort=max 83.71
    Printed as 83.71%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-6-luna; max_output_tokens=128000; reasoning_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 40 of 105 (GPT-6 Luna), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["openai/gpt-6-luna"].
    40 | GPT-6 Luna | 83.71%±1.95 | $0.1/$0.5 | 2m50s
    Every result from this document
  3. 38MedCode (Vals AI) model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000 44.69
    Printed as 44.69%Official leaderboard, measured Sep 2026Configuration: model ID openai/gpt-6-luna; reasoning_effort=max; max_output_tokens=128000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 38 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["openai/gpt-6-luna"].
    38 | GPT-6 Luna | 44.69%±2.30 | $0.1/$0.5 | 108.77s
    Every result from this document

Other OpenAI models: GPT-4.1, GPT-4.1 mini, GPT-4o, GPT-5, GPT-5.1, GPT-5.2, GPT-5.3-Codex, GPT-5.4, GPT-5.4 mini, GPT-5.4 nano, GPT-5.5, GPT-5.5 Instant, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5 mini, GPT-5 nano, GPT-6.1 Sol, GPT-6 Astra, GPT-6 Sol, GPT OSS 120B, GPT OSS 20B, o3, o4-mini