HealthBench Professional
Physician-selected workplace tasks, from care consults to documentation and research, graded on physician-written rubrics.
More about this boardLess
Physician-selected workplace tasks, from care consults to documentation and research, graded on physician-written rubrics.
525 tasks that physicians picked out of 15,079 real workplace AI conversations, spanning care consults, clinical documentation, and medical research, each judged on a rubric physicians wrote for it.
Grader GPT-5.4 at low reasoning effort with a length adjustment. Physician-written responses score 0.437 on the same rubrics.
Published by OpenAI, released 22 Apr 2026. 525 physician-authored tasks. 0 to 1, higher is better.
- Length-adjusted rubric score × 100. April 2026 paper, overall benchmark; GPT-5.4 at highest reasoning; eight samples per task. Unit: points.
- Labs on the board
- 3
- Scale
- 0 to 1
- Tasks
- 525
- Grader
- GPT-5.4, low reasoning effort
- Use cases
- 3
- Countries
- 50
- Languages
- 52
- Physicians involved
- 190
- Pricing note
- GPT-5.6 models charge higher rates above 272K input tokens. MAI-Thinking-1 is in public preview on Microsoft Foundry without final list pricing.
- Specialties
- 26
- Paper version
- April 2026 paper v1; length-adjusted primary score
- Candidate conversations screened
- 15,079
- Physician-written responses score
- 0.437
- Rows
- 31
- Models
- 26
- Labs
- 5
- Last measured
- Sep 2026
- Leader
- 0.703GPT-6 Astra (Anthropic run)
Ranking
0 to 1, higher is better
- Anthropic
- OpenAI
- Meta
- Other labs
20 of 31 rows
Rows and sources
Open a row for the quote, the page and the document.
1GPT-6 Astra (Anthropic run) Anthropic public-API reproduction; max effort; no system prompt; Opus 4.8 grade… 0.703
Printed as 70.3%Independent run, measured Sep 2026Configuration: Anthropic public-API reproduction; max effort; no system prompt; Opus 4.8 grader; length-adjusted; raw 74.0%Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, paragraph above Figure 8.15.B, HealthBench Professional after length adjustment.On HealthBench Professional at max effort, length adjustment changes the ranking. After: GPT-6 Astra (70.3%) > Claude Sonnet 5.5 (69.2%) > Claude Opus 5.5 (65.6%) > Claude Fable 5.1 (62.1%).
Every result from this document2Claude Sonnet 5.5 Anthropic; max effort; Opus 4.8 grader; safety classifiers enabled; length-adju… 0.692
Printed as 69.2%Vendor-reported, measured Sep 2026Configuration: Anthropic; max effort; Opus 4.8 grader; safety classifiers enabled; length-adjusted; raw 77.1%Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Section 8.15.2 and paragraph above Figure 8.15.B.On HealthBench Professional at max effort, length adjustment changes the ranking. After: GPT-6 Astra (70.3%) > Claude Sonnet 5.5 (69.2%) > Claude Opus 5.5 (65.6%) > Claude Fable 5.1 (62.1%).
Every result from this document3Claude Fable 5 length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Op… 0.660
Printed as 66.0Vendor-reported, measured Jun 2026Configuration: length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 70.3%). Measured as Claude Mythos 5; Anthropic itself prints 66.0 in the Fable 5 column of the Opus 5 card with footnote Mythos 5.Claude Fable 5 and Claude Mythos 5 System Card system card, Anthropic, 9 Jun 2026. p. 252, Table 8.1.A, row HealthBench Professional, column Mythos 5 (Fable 5 column '-'); Figure 8.18.2.A p. 298 bar label 66.0% on Claude Mythos 5HealthBench Professional 66.0 - 64.7 56.9 51.8 -
Every result from this document- Also printed in System Card: Claude Opus 5: 66.0, p. 152, Table 8.1.A, column Fable 5 (footnote 11: Mythos 5)
- A different number in Claude Fable 5.1 and Claude Mythos 5.1 System Card: 63.3%, p. 167, Table 8.1.A, column 'Claude Fable 5/Mythos 5'
4Claude Opus 5.5 Anthropic; adaptive max effort; Opus 4.8 grader; five trials; no tools or custo… 0.656
Printed as 65.6%Vendor-reported, measured Sep 2026Configuration: Anthropic; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt; safety classifiers and refusal fallback to Opus 5; length-adjusted; raw 77.1%Claude Opus 5.5 System Card system card, Anthropic, 22 Sep 2026. pp. 213–214, Section 8.15.2, paragraph and Figure 8.15.2.A.8.15.2 HealthBench Professional results length adjustment, which penalizes verbose model responses, Claude Opus 5.5 achieved a score of 65.6%.
Every result from this document5GPT-6 Astra length-adjusted, max reasoning effort (69.5 unadjusted, 4,097 mean response cha… 0.647
Printed as 64.7Vendor-reported, measured Sep 2026Configuration: length-adjusted, max reasoning effort (69.5 unadjusted, 4,097 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'.GPT-6 Astra System Card - HealthBench (Deployment Safety Hub) system card, OpenAI, 3 Sep 2026. Section 11.4.1, Table 29, HealthBench Professional length-adjusted, GPT-6 Astra column; September 22 correction.Evaluation | GPT-5.5 | GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna HealthBench Professional length-adjusted | 51.8 (57.2, 3818) | 60.5 (64.1, 3228) | 57.7 (62.4, 3618) | 55.7 (59.8, 3389) | 64.7 (68.2, 3185) | 60.8 (59.5, 1573) | 60.8 (61.2, 2119)
Every result from this document- Also printed in GPT-6 Astra: A new generation of intelligence: 63.4%, Science and Health table
- Mirrored by GPT-6 Astra System Card - HealthBench (Deployment Safety Hub): 63.4, section 6.1, HTML rendering of the same card
- A different number in GPT-6 Astra System Card: 63.4, p. 19, sec. 6.1 (Table 6 cell, column 'gpt-6 Astra': 63.4 (69.5, 4097))
6Claude Fable 5 (September card) Measured as Claude Fable 5; Anthropic September card; length-adjusted; adaptive… 0.633
Printed as 63.3%Vendor-reported, measured Sep 2026Configuration: Measured as Claude Fable 5; Anthropic September card; length-adjusted; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt; raw 68.9%Claude Fable 5.1 and Claude Mythos 5.1 System Card system card, Anthropic, 1 Sep 2026. p. 199, Figure 8.17.2.A, Claude Fable 5 length-adjusted bar (visually read printed labels).HealthBench Professional | Length-adjusted score | Claude Fable 5 | 63.3%
Every result from this document7Claude Fable 5.1 length-adjusted (method published in the HealthBench Professional paper); Anthr… 0.621
Printed as 62.1%Vendor-reported, measured Sep 2026Configuration: length-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%).Claude Fable 5.1 and Claude Mythos 5.1 System Card system card, Anthropic, 1 Sep 2026. p. 199, sec. 8.17.2 (same figure printed as 62.1% in Table 8.1.A, p. 167)On HealthBench Professional, Claude Fable 5.1 achieved a raw score of 74.2%, ahead of Claude Opus 5 at 73.4%, Fable 5 at 68.9%, and Claude Sonnet 5 at 62.4%. After length adjustment, which penalizes verbose model responses, Fable 5.1 achieved a score of 62.1%.
Every result from this document8GPT-6 Luna OpenAI; length-adjusted; maximum reasoning effort; 60.8 (61.2, 2119) (adjusted,… 0.608
Printed as 60.8Vendor-reported, measured Sep 2026Configuration: OpenAI; length-adjusted; maximum reasoning effort; 60.8 (61.2, 2119) (adjusted, raw, mean response characters)GPT-6 Astra System Card - HealthBench (Deployment Safety Hub) system card, OpenAI, 3 Sep 2026. Section 11.4.1, Table 29, HealthBench Professional length-adjusted, GPT-6 Luna column; September 22 correction.Evaluation | GPT-5.5 | GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna HealthBench Professional length-adjusted | 51.8 (57.2, 3818) | 60.5 (64.1, 3228) | 57.7 (62.4, 3618) | 55.7 (59.8, 3389) | 64.7 (68.2, 3185) | 60.8 (59.5, 1573) | 60.8 (61.2, 2119)
Every result from this document8GPT-6 Sol OpenAI; length-adjusted; maximum reasoning effort; 60.8 (59.5, 1573) (adjusted,… 0.608
Printed as 60.8Vendor-reported, measured Sep 2026Configuration: OpenAI; length-adjusted; maximum reasoning effort; 60.8 (59.5, 1573) (adjusted, raw, mean response characters)GPT-6 Astra System Card - HealthBench (Deployment Safety Hub) system card, OpenAI, 3 Sep 2026. Section 11.4.1, Table 29, HealthBench Professional length-adjusted, GPT-6 Sol column; September 22 correction.Evaluation | GPT-5.5 | GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna HealthBench Professional length-adjusted | 51.8 (57.2, 3818) | 60.5 (64.1, 3228) | 57.7 (62.4, 3618) | 55.7 (59.8, 3389) | 64.7 (68.2, 3185) | 60.8 (59.5, 1573) | 60.8 (61.2, 2119)
Every result from this document10GPT-5.6 Sol length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response cha… 0.605
Printed as 60.5Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492.GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOLHealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
Every result from this document- Also printed in GPT-5.6 Preview System Card: 60.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL
11Claude Opus 5 length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Op… 0.598
Printed as 59.8%Vendor-reported, measured Jul 2026Configuration: length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%).System Card: Claude Opus 5 system card, Anthropic, 24 Jul 2026. p. 189, section 8.15.2 HealthBench Professional results; also Table 8.1.A p. 152 ('HealthBench Professional 59.8 ...')Claude Opus 5 achieved a raw score of 73.4%, which is the highest amongst all Claude models, ahead of Claude Mythos 5 at 70.3%, Claude Opus 4.8 at 60.3%, and Claude Sonnet 5 at 62.4%. After length adjustment, which penalizes verbose model responses, Claude Opus 5 achieved a score of 59.8%.
Every result from this document- Also printed in Claude Fable 5.1 and Claude Mythos 5.1 System Card: 59.8%, p. 167, Table 8.1.A, column 'Claude Opus 5'; raw 73.4% also reprinted on p. 199, sec. 8.17.2
12Muse Spark 1.1 length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model… 0.593
Printed as 59.3Vendor-reported, measured Jul 2026Configuration: length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44)Muse Spark 1.1 Evaluation Report model card, Meta, 9 Jul 2026. p. 101, Figure 44 'General capability benchmark results' (image), row HealthBench Professional, column Muse Spark 1.1; protocol p. 104 (printed 103): HealthBench Pro comprises 525 evaluation data points graded by rubrics. We use GPT-5.4 with low reasoning effort as the grader and report the length-normalized rubric score as done in their paper.Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
Every result from this document13Claude Sonnet 5 length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Op… 0.578
Printed as 57.8Vendor-reported, measured Jun 2026Configuration: length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).System Card: Claude Sonnet 5 system card, Anthropic, 30 Jun 2026. p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 5; Figure 8.12.2.A p. 139HealthBench Professional 57.8 44.2 51.8 -
Every result from this document- Also printed in System Card: Claude Opus 5: 57.8%, p. 189, section 8.15.2, Figure 8.15.2.A
14GPT-5.6 Terra length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response cha… 0.577
Printed as 57.7Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRAHealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
Every result from this document- Also printed in GPT-5.6 Preview System Card: 57.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA
15Claude Opus 4.8 (Opus 4.8 grader) Anthropic June evaluation; length-adjusted; adaptive max effort; Opus 4.8 grade… 0.574
Printed as 57.4%Vendor-reported, measured Jun 2026Configuration: Anthropic June evaluation; length-adjusted; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system promptSystem Card: Claude Sonnet 5 system card, Anthropic, 30 Jun 2026. p. 139, Figure 8.12.2.A, Claude Opus 4.8 bar (visually read printed labels).HealthBench Professional | Length-adjusted score (%) | Claude Opus 4.8 | 57.4%
Every result from this document16Grok 4.7 SpaceXAI evaluation; xHigh effort; HealthBench Professional 0.567
Printed as 56.7%Vendor-reported, measured Sep 2026Configuration: SpaceXAI evaluation; xHigh effort; HealthBench ProfessionalIntroducing Grok 4.7 launch post, SpaceXAI, 21 Sep 2026. Model Improvements comparison table, Clinical reasoning / HealthBench Professional row; Grok 4.7 column; September 21, 2026.Grok 4.7 xHigh Grok 4.6 High GPT-5.6 Sol Max Fable 5.1 Max Clinical reasoningHealthBench Professional 56.7% 48.5% 60.5% 62.1%
Every result from this document17Claude Opus 4.8 length-adjusted, adaptive thinking at max effort, Claude Sonnet 4.6 grader (the… 0.558
Printed as 55.8%Vendor-reported, measured May 2026Configuration: length-adjusted, adaptive thinking at max effort, Claude Sonnet 4.6 grader (the grader used by the Opus 4.8 card itself); the other Claude rows on this board use the Claude Opus 4.8 grader, under which Anthropic later prints 57.4 for this model.System Card: Claude Opus 4.8 system card, Anthropic, 28 May 2026. p. 228, section 8.14.1 HealthBench Professional; Figure 8.14.A p. 229Claude Opus 4.8 scores 55.8%, a meaningful improvement over Claude Opus 4.7 at 51.9% and Claude Sonnet 4.6 at 41.7%.
Every result from this document- A different number in Claude Fable 5 and Claude Mythos 5 System Card: 56.9, p. 252, Table 8.1.A, column Opus 4.8
- A different number in System Card: Claude Opus 5: 57.4, p. 152, Table 8.1.A, column Opus 4.8
18GPT-5.6 Luna length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response cha… 0.557
Printed as 55.7Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493.GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNAHealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
Every result from this document- Also printed in GPT-5.6 Preview System Card: 55.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA
19Muse Spark length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the… 0.541
Printed as 54.1Vendor-reported, measured Jul 2026Configuration: length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44Muse Spark 1.1 Evaluation Report model card, Meta, 9 Jul 2026. p. 101, Figure 44 (image), row HealthBench Professional, column Muse Spark; protocol p. 104 (printed 103)Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
Every result from this document20GPT-5.6 Sol (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 Augus… 0.540
Printed as 54.0Vendor-reported, measured Aug 2026Configuration: ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars)GPT-5.6 - August Updates (system card addendum) system card, OpenAI, 6 Aug 2026. p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)HealthBench Professional 32.9 (33.8, 2,285) 38.4 (40.7, 2,775) 54.0 (56.6, 2,894) 44.1 (46.8, 2,920)
Every result from this document21Claude Opus 4.7 length-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 t… 0.519
Printed as 51.9%Vendor-reported, measured May 2026Configuration: length-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 trials; comparison model in the Opus 4.8 cardSystem Card: Claude Opus 4.8 system card, Anthropic, 28 May 2026. p. 228, section 8.14.1 HealthBench Professional; Figure 8.14.A p. 229Claude Opus 4.8 scores 55.8%, a meaningful improvement over Claude Opus 4.7 at 51.9% and Claude Sonnet 4.6 at 41.7%.
Every result from this document22GPT-5.5 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5… 0.518
Printed as 51.8Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (57.2 unadjusted, 3818 chars)GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
Every result from this document- Also printed in GPT-5.6 Preview System Card: 51.8, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.5
23Grok 4.6 SpaceXAI evaluation; High effort; HealthBench Professional 0.485
Printed as 48.5%Vendor-reported, measured Sep 2026Configuration: SpaceXAI evaluation; High effort; HealthBench ProfessionalIntroducing Grok 4.7 launch post, SpaceXAI, 21 Sep 2026. Model Improvements comparison table, Clinical reasoning / HealthBench Professional row; Grok 4.6 column; September 21, 2026.Grok 4.7 xHigh Grok 4.6 High GPT-5.6 Sol Max Fable 5.1 Max Clinical reasoningHealthBench Professional 56.7% 48.5% 60.5% 62.1%
Every result from this document24GPT-5.4 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5… 0.481
Printed as 48.1Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars)GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.4HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
Every result from this document- Also printed in GPT-5.6 Preview System Card: 48.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.4
25GPT-5 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5… 0.462
Printed as 46.2Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars)GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
Every result from this document- Also printed in GPT-5.6 Preview System Card: 46.2, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5
26GPT-5.2 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5… 0.459
Printed as 45.9Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars)GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
Every result from this document- Also printed in GPT-5.6 Preview System Card: 45.9, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.2
27Claude Sonnet 4.6 length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Op… 0.442
Printed as 44.2Vendor-reported, measured Jun 2026Configuration: length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Sonnet 5 card. Other Anthropic prints: 44.4% (Fable card Figure 8.18.2.A), 41.7% (Opus 4.8 card, Sonnet 4.6 grader)System Card: Claude Sonnet 5 system card, Anthropic, 30 Jun 2026. p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 4.6HealthBench Professional 57.8 44.2 51.8 -
Every result from this document- A different number in System Card: Claude Opus 4.8: 41.7%, p. 228, section 8.14.1
28GPT-5.6 Luna (August) ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 Augus… 0.441
Printed as 44.1Vendor-reported, measured Aug 2026Configuration: ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars)GPT-5.6 - August Updates (system card addendum) system card, OpenAI, 6 Aug 2026. p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)HealthBench Professional 32.9 (33.8, 2,285) 38.4 (40.7, 2,775) 54.0 (56.6, 2,894) 44.1 (46.8, 2,920)
Every result from this document29GPT-5.1 length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5… 0.396
Printed as 39.6Vendor-reported, measured Jun 2026Configuration: length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars)GPT-5.6 System Card system card, OpenAI, 9 Jul 2026. Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.1HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
Every result from this document- Also printed in GPT-5.6 Preview System Card: 39.6, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.1
30GPT-5.5 Instant length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant s… 0.384
Printed as 38.4Vendor-reported, measured May 2026Configuration: length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.GPT-5.5 Instant System Card system card, OpenAI, 5 May 2026. Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANTHealthBench Professional 37.6 (40.4, 2,973) 35.7 (38.3, 2,872) 32.9 (33.8, 2,285) 38.4 (40.7, 2,775)
Every result from this document- Also printed in GPT-5.6 - August Updates (system card addendum): 38.4, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant
31MAI-Thinking-1 length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 gra… 0.350
Printed as 35Vendor-reported, measured Aug 2026Configuration: length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision.MAI-Thinking-1: Building a Hill-Climbing Machine model card, Microsoft AI, 12 Aug 2026. p. 54, Table 12 'Post-trained model evaluation results on various public benchmarks', Health group, column HealthBench Prof.; protocol Appendix K.6 p. 106: HealthBench Professional introduces a length penalty for the primary metric, to correct for a well-observed correlation between lengthy responses and artificially increased LLM-grader scores. For all reported scores, we use the standard GPT-5.4 grader and rubrics provided by OpenAI.Model AIR-Bench CyberSec Instruct CyberSec Auto Long Fact Truthful QA HealthBench Prof. MedXpert QA MAI-Thinking-1 88 63 63 98 88 35 43 Sonnet 4.6 88 62 56 98 88 38 49
Every result from this document
Documents
17
- Claude Fable 5 and Claude Mythos 5 System Cardsystem card, Anthropic, 9 Jun 2026Results it supports
- Claude Fable 5.1 and Claude Mythos 5.1 System Cardsystem card, Anthropic, 1 Sep 2026Results it supports
- Claude Opus 5.5 System Cardsystem card, Anthropic, 22 Sep 2026Results it supports
- Claude Sonnet 5.5 System Cardsystem card, Anthropic, 28 Sep 2026Results it supports
- GPT-5.5 Instant System Cardsystem card, OpenAI, 5 May 2026Results it supports
- GPT-5.6 - August Updates (system card addendum)system card, OpenAI, 6 Aug 2026Results it supports
- GPT-5.6 Preview System Cardsystem card, OpenAI, 26 Jun 2026Results it supports
- GPT-5.6 System Cardsystem card, OpenAI, 9 Jul 2026Results it supports
- GPT-6 Astra System Cardsystem card, OpenAI, 3 Sep 2026Results it supports
- GPT-6 Astra System Card - HealthBench (Deployment Safety Hub)system card, OpenAI, 3 Sep 2026Results it supports
- GPT-6 Astra: A new generation of intelligencelaunch post, OpenAI, 3 Sep 2026Results it supports
- Introducing Grok 4.7launch post, SpaceXAI, 21 Sep 2026Results it supports
- MAI-Thinking-1: Building a Hill-Climbing Machinemodel card, Microsoft AI, 12 Aug 2026Results it supports
- Muse Spark 1.1 Evaluation Reportmodel card, Meta, 9 Jul 2026Results it supports
- System Card: Claude Opus 4.8system card, Anthropic, 28 May 2026Results it supports
- System Card: Claude Opus 5system card, Anthropic, 24 Jul 2026Results it supports
- System Card: Claude Sonnet 5system card, Anthropic, 30 Jun 2026Results it supports