Clinical Benchmarks

HealthBench Professional

Physician-selected workplace tasks, from care consults to documentation and research, graded on physician-written rubrics.

OpenAIDedicated boardPaper
More about this boardLess

Physician-selected workplace tasks, from care consults to documentation and research, graded on physician-written rubrics.

525 tasks that physicians picked out of 15,079 real workplace AI conversations, spanning care consults, clinical documentation, and medical research, each judged on a rubric physicians wrote for it.

Grader GPT-5.4 at low reasoning effort with a length adjustment. Physician-written responses score 0.437 on the same rubrics.

Published by OpenAI, released 22 Apr 2026. 525 physician-authored tasks. 0 to 1, higher is better.

  • Length-adjusted rubric score × 100. April 2026 paper, overall benchmark; GPT-5.4 at highest reasoning; eight samples per task. Unit: points.
Labs on the board
3
Scale
0 to 1
Tasks
525
Grader
GPT-5.4, low reasoning effort
Use cases
3
Countries
50
Languages
52
Physicians involved
190
Pricing note
GPT-5.6 models charge higher rates above 272K input tokens. MAI-Thinking-1 is in public preview on Microsoft Foundry without final list pricing.
Specialties
26
Paper version
April 2026 paper v1; length-adjusted primary score
Candidate conversations screened
15,079
Physician-written responses score
0.437
Rows
31
Models
26
Labs
5
Last measured
Sep 2026
Leader
0.703GPT-6 Astra (Anthropic run)

Ranking

0 to 1, higher is better

  • Anthropic
  • OpenAI
  • Meta
  • Other labs
  1. 1GPT-6 Astra (Anthropic run)Anthropic public-API reproduction0.703
  2. 2Claude Sonnet 5.5Anthropic0.692
  3. 3Claude Fable 5length-adjusted0.660
  4. 4Claude Opus 5.5Anthropic0.656
  5. 5GPT-6 Astralength-adjusted0.647
  6. 6Claude Fable 5 (September card)Measured as Claude Fable 50.633
  7. 7Claude Fable 5.1length-adjusted0.621
  8. 8GPT-6 LunaOpenAI0.608
  9. 8GPT-6 SolOpenAI0.608
  10. 10GPT-5.6 Sollength-adjusted0.605
  11. 11Claude Opus 5length-adjusted0.598
  12. 12Muse Spark 1.1length-normalized0.593
  13. 13Claude Sonnet 5length-adjusted0.578
  14. 14GPT-5.6 Terralength-adjusted0.577
  15. 15Claude Opus 4.8 (Opus 4.8 grader)Anthropic June evaluation0.574
  16. 16Grok 4.7SpaceXAI evaluation0.567
  17. 17Claude Opus 4.8length-adjusted0.558
  18. 18GPT-5.6 Lunalength-adjusted0.557
  19. 19Muse Sparklength-normalized0.541
  20. 20GPT-5.6 Sol (August)ChatGPT production0.540

20 of 31 rows

Rows and sources

Open a row for the quote, the page and the document.

  1. 1GPT-6 Astra (Anthropic run) Anthropic public-API reproduction; max effort; no system prompt; Opus 4.8 grade… 0.703
    Printed as 70.3%Independent run, measured Sep 2026Configuration: Anthropic public-API reproduction; max effort; no system prompt; Opus 4.8 grader; length-adjusted; raw 74.0%
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, paragraph above Figure 8.15.B, HealthBench Professional after length adjustment.
    On HealthBench Professional at max effort, length adjustment changes the ranking. After: GPT-6 Astra (70.3%) > Claude Sonnet 5.5 (69.2%) > Claude Opus 5.5 (65.6%) > Claude Fable 5.1 (62.1%).
    Every result from this document
  2. 2Claude Sonnet 5.5 Anthropic; max effort; Opus 4.8 grader; safety classifiers enabled; length-adju… 0.692
    Printed as 69.2%Vendor-reported, measured Sep 2026Configuration: Anthropic; max effort; Opus 4.8 grader; safety classifiers enabled; length-adjusted; raw 77.1%
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Section 8.15.2 and paragraph above Figure 8.15.B.
    On HealthBench Professional at max effort, length adjustment changes the ranking. After: GPT-6 Astra (70.3%) > Claude Sonnet 5.5 (69.2%) > Claude Opus 5.5 (65.6%) > Claude Fable 5.1 (62.1%).
    Every result from this document
  3. 3Claude Fable 5 length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Op… 0.660
    Printed as 66.0Vendor-reported, measured Jun 2026Configuration: length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 70.3%). Measured as Claude Mythos 5; Anthropic itself prints 66.0 in the Fable 5 column of the Opus 5 card with footnote Mythos 5.
    Claude Fable 5 and Claude Mythos 5 System Card system card, Anthropic, 9 Jun 2026. p. 252, Table 8.1.A, row HealthBench Professional, column Mythos 5 (Fable 5 column '-'); Figure 8.18.2.A p. 298 bar label 66.0% on Claude Mythos 5
    HealthBench Professional 66.0 - 64.7 56.9 51.8 -
    Every result from this document
  4. 4Claude Opus 5.5 Anthropic; adaptive max effort; Opus 4.8 grader; five trials; no tools or custo… 0.656
    Printed as 65.6%Vendor-reported, measured Sep 2026Configuration: Anthropic; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt; safety classifiers and refusal fallback to Opus 5; length-adjusted; raw 77.1%
    Claude Opus 5.5 System Card system card, Anthropic, 22 Sep 2026. pp. 213–214, Section 8.15.2, paragraph and Figure 8.15.2.A.
    8.15.2 HealthBench Professional results length adjustment, which penalizes verbose model responses, Claude Opus 5.5 achieved a score of 65.6%.
    Every result from this document
  5. 5GPT-6 Astra length-adjusted, max reasoning effort (69.5 unadjusted, 4,097 mean response cha… 0.647
    Printed as 64.7Vendor-reported, measured Sep 2026Configuration: length-adjusted, max reasoning effort (69.5 unadjusted, 4,097 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'.
    GPT-6 Astra System Card - HealthBench (Deployment Safety Hub) system card, OpenAI, 3 Sep 2026. Section 11.4.1, Table 29, HealthBench Professional length-adjusted, GPT-6 Astra column; September 22 correction.
    Evaluation | GPT-5.5 | GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna HealthBench Professional length-adjusted | 51.8 (57.2, 3818) | 60.5 (64.1, 3228) | 57.7 (62.4, 3618) | 55.7 (59.8, 3389) | 64.7 (68.2, 3185) | 60.8 (59.5, 1573) | 60.8 (61.2, 2119)
    Every result from this document
  6. 6Claude Fable 5 (September card) Measured as Claude Fable 5; Anthropic September card; length-adjusted; adaptive… 0.633
    Printed as 63.3%Vendor-reported, measured Sep 2026Configuration: Measured as Claude Fable 5; Anthropic September card; length-adjusted; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt; raw 68.9%
    Claude Fable 5.1 and Claude Mythos 5.1 System Card system card, Anthropic, 1 Sep 2026. p. 199, Figure 8.17.2.A, Claude Fable 5 length-adjusted bar (visually read printed labels).
    HealthBench Professional | Length-adjusted score | Claude Fable 5 | 63.3%
    Every result from this document
  7. 7Claude Fable 5.1 length-adjusted (method published in the HealthBench Professional paper); Anthr… 0.621
    Printed as 62.1%Vendor-reported, measured Sep 2026Configuration: length-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%).
    Claude Fable 5.1 and Claude Mythos 5.1 System Card system card, Anthropic, 1 Sep 2026. p. 199, sec. 8.17.2 (same figure printed as 62.1% in Table 8.1.A, p. 167)
    On HealthBench Professional, Claude Fable 5.1 achieved a raw score of 74.2%, ahead of Claude Opus 5 at 73.4%, Fable 5 at 68.9%, and Claude Sonnet 5 at 62.4%. After length adjustment, which penalizes verbose model responses, Fable 5.1 achieved a score of 62.1%.
    Every result from this document
  8. 8GPT-6 Luna OpenAI; length-adjusted; maximum reasoning effort; 60.8 (61.2, 2119) (adjusted,… 0.608
    Printed as 60.8Vendor-reported, measured Sep 2026Configuration: OpenAI; length-adjusted; maximum reasoning effort; 60.8 (61.2, 2119) (adjusted, raw, mean response characters)
    GPT-6 Astra System Card - HealthBench (Deployment Safety Hub) system card, OpenAI, 3 Sep 2026. Section 11.4.1, Table 29, HealthBench Professional length-adjusted, GPT-6 Luna column; September 22 correction.
    Evaluation | GPT-5.5 | GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna HealthBench Professional length-adjusted | 51.8 (57.2, 3818) | 60.5 (64.1, 3228) | 57.7 (62.4, 3618) | 55.7 (59.8, 3389) | 64.7 (68.2, 3185) | 60.8 (59.5, 1573) | 60.8 (61.2, 2119)
    Every result from this document
  9. 8GPT-6 Sol OpenAI; length-adjusted; maximum reasoning effort; 60.8 (59.5, 1573) (adjusted,… 0.608
    Printed as 60.8Vendor-reported, measured Sep 2026Configuration: OpenAI; length-adjusted; maximum reasoning effort; 60.8 (59.5, 1573) (adjusted, raw, mean response characters)
    GPT-6 Astra System Card - HealthBench (Deployment Safety Hub) system card, OpenAI, 3 Sep 2026. Section 11.4.1, Table 29, HealthBench Professional length-adjusted, GPT-6 Sol column; September 22 correction.
    Evaluation | GPT-5.5 | GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna HealthBench Professional length-adjusted | 51.8 (57.2, 3818) | 60.5 (64.1, 3228) | 57.7 (62.4, 3618) | 55.7 (59.8, 3389) | 64.7 (68.2, 3185) | 60.8 (59.5, 1573) | 60.8 (61.2, 2119)
    Every result from this document
  10. 10GPT-5.6 Sol length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response cha… 0.605
    Printed as 60.5Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492.
    GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Every result from this document
    • Also printed in GPT-5.6 Preview System Card: 60.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL
  11. 11Claude Opus 5 length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Op… 0.598
    Printed as 59.8%Vendor-reported, measured Jul 2026Configuration: length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%).
    System Card: Claude Opus 5 system card, Anthropic, 24 Jul 2026. p. 189, section 8.15.2 HealthBench Professional results; also Table 8.1.A p. 152 ('HealthBench Professional 59.8 ...')
    Claude Opus 5 achieved a raw score of 73.4%, which is the highest amongst all Claude models, ahead of Claude Mythos 5 at 70.3%, Claude Opus 4.8 at 60.3%, and Claude Sonnet 5 at 62.4%. After length adjustment, which penalizes verbose model responses, Claude Opus 5 achieved a score of 59.8%.
    Every result from this document
  12. 12Muse Spark 1.1 length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model… 0.593
    Printed as 59.3Vendor-reported, measured Jul 2026Configuration: length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44)
    Muse Spark 1.1 Evaluation Report model card, Meta, 9 Jul 2026. p. 101, Figure 44 'General capability benchmark results' (image), row HealthBench Professional, column Muse Spark 1.1; protocol p. 104 (printed 103): HealthBench Pro comprises 525 evaluation data points graded by rubrics. We use GPT-5.4 with low reasoning effort as the grader and report the length-normalized rubric score as done in their paper.
    Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
    Every result from this document
  13. 13Claude Sonnet 5 length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Op… 0.578
    Printed as 57.8Vendor-reported, measured Jun 2026Configuration: length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).
    System Card: Claude Sonnet 5 system card, Anthropic, 30 Jun 2026. p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 5; Figure 8.12.2.A p. 139
    HealthBench Professional 57.8 44.2 51.8 -
    Every result from this document
  14. 14GPT-5.6 Terra length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response cha… 0.577
    Printed as 57.7Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
    GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Every result from this document
    • Also printed in GPT-5.6 Preview System Card: 57.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA
  15. 15Claude Opus 4.8 (Opus 4.8 grader) Anthropic June evaluation; length-adjusted; adaptive max effort; Opus 4.8 grade… 0.574
    Printed as 57.4%Vendor-reported, measured Jun 2026Configuration: Anthropic June evaluation; length-adjusted; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt
    System Card: Claude Sonnet 5 system card, Anthropic, 30 Jun 2026. p. 139, Figure 8.12.2.A, Claude Opus 4.8 bar (visually read printed labels).
    HealthBench Professional | Length-adjusted score (%) | Claude Opus 4.8 | 57.4%
    Every result from this document
  16. 16Grok 4.7 SpaceXAI evaluation; xHigh effort; HealthBench Professional 0.567
    Printed as 56.7%Vendor-reported, measured Sep 2026Configuration: SpaceXAI evaluation; xHigh effort; HealthBench Professional
    Introducing Grok 4.7 launch post, SpaceXAI, 21 Sep 2026. Model Improvements comparison table, Clinical reasoning / HealthBench Professional row; Grok 4.7 column; September 21, 2026.
    Grok 4.7 xHigh Grok 4.6 High GPT-5.6 Sol Max Fable 5.1 Max Clinical reasoningHealthBench Professional 56.7% 48.5% 60.5% 62.1%
    Every result from this document
  17. 17Claude Opus 4.8 length-adjusted, adaptive thinking at max effort, Claude Sonnet 4.6 grader (the… 0.558
    Printed as 55.8%Vendor-reported, measured May 2026Configuration: length-adjusted, adaptive thinking at max effort, Claude Sonnet 4.6 grader (the grader used by the Opus 4.8 card itself); the other Claude rows on this board use the Claude Opus 4.8 grader, under which Anthropic later prints 57.4 for this model.
    System Card: Claude Opus 4.8 system card, Anthropic, 28 May 2026. p. 228, section 8.14.1 HealthBench Professional; Figure 8.14.A p. 229
    Claude Opus 4.8 scores 55.8%, a meaningful improvement over Claude Opus 4.7 at 51.9% and Claude Sonnet 4.6 at 41.7%.
    Every result from this document
  18. 18GPT-5.6 Luna length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response cha… 0.557
    Printed as 55.7Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493.
    GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Every result from this document
    • Also printed in GPT-5.6 Preview System Card: 55.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA
  19. 19Muse Spark length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the… 0.541
    Printed as 54.1Vendor-reported, measured Jul 2026Configuration: length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44
    Muse Spark 1.1 Evaluation Report model card, Meta, 9 Jul 2026. p. 101, Figure 44 (image), row HealthBench Professional, column Muse Spark; protocol p. 104 (printed 103)
    Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
    Every result from this document
  20. 20GPT-5.6 Sol (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 Augus… 0.540
    Printed as 54.0Vendor-reported, measured Aug 2026Configuration: ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars)
    GPT-5.6 - August Updates (system card addendum) system card, OpenAI, 6 Aug 2026. p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
    HealthBench Professional 32.9 (33.8, 2,285) 38.4 (40.7, 2,775) 54.0 (56.6, 2,894) 44.1 (46.8, 2,920)
    Every result from this document
  21. 21Claude Opus 4.7 length-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 t… 0.519
    Printed as 51.9%Vendor-reported, measured May 2026Configuration: length-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 trials; comparison model in the Opus 4.8 card
    System Card: Claude Opus 4.8 system card, Anthropic, 28 May 2026. p. 228, section 8.14.1 HealthBench Professional; Figure 8.14.A p. 229
    Claude Opus 4.8 scores 55.8%, a meaningful improvement over Claude Opus 4.7 at 51.9% and Claude Sonnet 4.6 at 41.7%.
    Every result from this document
  22. 22GPT-5.5 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5… 0.518
    Printed as 51.8Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (57.2 unadjusted, 3818 chars)
    GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Every result from this document
    • Also printed in GPT-5.6 Preview System Card: 51.8, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.5
  23. 23Grok 4.6 SpaceXAI evaluation; High effort; HealthBench Professional 0.485
    Printed as 48.5%Vendor-reported, measured Sep 2026Configuration: SpaceXAI evaluation; High effort; HealthBench Professional
    Introducing Grok 4.7 launch post, SpaceXAI, 21 Sep 2026. Model Improvements comparison table, Clinical reasoning / HealthBench Professional row; Grok 4.6 column; September 21, 2026.
    Grok 4.7 xHigh Grok 4.6 High GPT-5.6 Sol Max Fable 5.1 Max Clinical reasoningHealthBench Professional 56.7% 48.5% 60.5% 62.1%
    Every result from this document
  24. 24GPT-5.4 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5… 0.481
    Printed as 48.1Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars)
    GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.4
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Every result from this document
    • Also printed in GPT-5.6 Preview System Card: 48.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.4
  25. 25GPT-5 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5… 0.462
    Printed as 46.2Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars)
    GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Every result from this document
    • Also printed in GPT-5.6 Preview System Card: 46.2, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5
  26. 26GPT-5.2 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5… 0.459
    Printed as 45.9Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars)
    GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Every result from this document
    • Also printed in GPT-5.6 Preview System Card: 45.9, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.2
  27. 27Claude Sonnet 4.6 length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Op… 0.442
    Printed as 44.2Vendor-reported, measured Jun 2026Configuration: length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Sonnet 5 card. Other Anthropic prints: 44.4% (Fable card Figure 8.18.2.A), 41.7% (Opus 4.8 card, Sonnet 4.6 grader)
    System Card: Claude Sonnet 5 system card, Anthropic, 30 Jun 2026. p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 4.6
    HealthBench Professional 57.8 44.2 51.8 -
    Every result from this document
  28. 28GPT-5.6 Luna (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 Augus… 0.441
    Printed as 44.1Vendor-reported, measured Aug 2026Configuration: ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars)
    GPT-5.6 - August Updates (system card addendum) system card, OpenAI, 6 Aug 2026. p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
    HealthBench Professional 32.9 (33.8, 2,285) 38.4 (40.7, 2,775) 54.0 (56.6, 2,894) 44.1 (46.8, 2,920)
    Every result from this document
  29. 29GPT-5.1 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5… 0.396
    Printed as 39.6Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars)
    GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.1
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Every result from this document
    • Also printed in GPT-5.6 Preview System Card: 39.6, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.1
  30. 30GPT-5.5 Instant length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant s… 0.384
    Printed as 38.4Vendor-reported, measured May 2026Configuration: length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
    GPT-5.5 Instant System Card system card, OpenAI, 5 May 2026. Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT
    HealthBench Professional 37.6 (40.4, 2,973) 35.7 (38.3, 2,872) 32.9 (33.8, 2,285) 38.4 (40.7, 2,775)
    Every result from this document
  31. 31MAI-Thinking-1 length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 gra… 0.350
    Printed as 35Vendor-reported, measured Aug 2026Configuration: length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision.
    MAI-Thinking-1: Building a Hill-Climbing Machine model card, Microsoft AI, 12 Aug 2026. p. 54, Table 12 'Post-trained model evaluation results on various public benchmarks', Health group, column HealthBench Prof.; protocol Appendix K.6 p. 106: HealthBench Professional introduces a length penalty for the primary metric, to correct for a well-observed correlation between lengthy responses and artificially increased LLM-grader scores. For all reported scores, we use the standard GPT-5.4 grader and rubrics provided by OpenAI.
    Model AIR-Bench CyberSec Instruct CyberSec Auto Long Fact Truthful QA HealthBench Prof. MedXpert QA MAI-Thinking-1 88 63 63 98 88 35 43 Sonnet 4.6 88 62 56 98 88 38 49
    Every result from this document

Documents

17