Clinical Benchmarks

Sources

218 of 218 rows documented · 52 documents, 40 first-party · updated September 8, 2026

Every number on this index was copied out of a document, and this page is the list of those documents. A row records where its score was read: the publication itself, the passage or table it sits in when that has been captured, who published it, when the index retrieved it, and how far the reading has been checked. First-party documents rank highest, meaning a vendor's system card for a vendor-reported number or the maintainers' own board for a leaderboard number; aggregator pages that repeat a figure are kept as corroboration and never stand in for the original. Rows whose document has not been located yet say so instead of disappearing.

Better citations do not make the boards comparable. Each benchmark below keeps its own scale, grader and task set, so a score on one section says nothing about a score on the next, and no number on this page should be lined up against a number from another section. The comparability rules are on the methodology page.

Reading the confidence field

HealthBench Professional

22 rows · 0 to 1

Board source: healthbenchprofessional.com · paper: arxiv.org · full board: healthbenchprofessional.com

  1. Claude Fable 5 Anthropic0.660as of 2026-06
    length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 70.3%). Measured as Claude Mythos 5; Anthropic itself prints 66.0 in the Fable 5 column of the Opus 5 card with footnote Mythos 5.
    publisher
    Anthropic
    locator
    p. 252, Table 8.1.A, row HealthBench Professional, column Mythos 5 (Fable 5 column '-'); Figure 8.18.2.A p. 298 bar label 66.0% on Claude Mythos 5
    reported by
    vendor-reported
    published
    2026-06-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional 66.0 - 64.7 56.9 51.8 -
    also reported in System Card: Claude Opus 5 (system card, 66.0, p. 152, Table 8.1.A, column Fable 5 (footnote 11: Mythos 5)); Claude Fable 5.1 and Claude Mythos 5.1 System Card (system card, 63.3%, conflicting, p. 167, Table 8.1.A, column 'Claude Fable 5/Mythos 5')
  2. GPT-6 Astra OpenAI0.634as of 2026-09
    length-adjusted, max reasoning effort (69.5 unadjusted, 4,097 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'.
    GPT-6 Astra System Card · system card · first-party
    publisher
    OpenAI
    locator
    p. 19, sec. 6.1 (Table 6 cell, column 'gpt-6 Astra': 63.4 (69.5, 4097))
    reported by
    vendor-reported
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    Astra has a length-adjusted HealthBench Professional score of 63.4 (+2.9 relative to GPT-5.6 Sol), HealthBench score of 58.1 (+1.1), HealthBench Hard score of 36.3 (+3.2), and HealthBench Consensus score of 95.8 (+0.3).
    also reported in GPT-6 Astra: A new generation of intelligence (launch post, 63.4%, Science and Health table); GPT-6 Astra System Card - HealthBench (Deployment Safety Hub) (system card, 63.4, mirror, section 6.1, HTML rendering of the same card)
  3. Claude Fable 5.1 Anthropic0.621as of 2026-09
    length-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%).
    publisher
    Anthropic
    locator
    p. 199, sec. 8.17.2 (same figure printed as 62.1% in Table 8.1.A, p. 167)
    reported by
    vendor-reported
    published
    2026-09-01
    retrieved
    2026-09-07
    confidence
    verified
    On HealthBench Professional, Claude Fable 5.1 achieved a raw score of 74.2%, ahead of Claude Opus 5 at 73.4%, Fable 5 at 68.9%, and Claude Sonnet 5 at 62.4%. After length adjustment, which penalizes verbose model responses, Fable 5.1 achieved a score of 62.1%.
  4. GPT-5.6 Sol OpenAI0.605as of 2026-06
    length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    also reported in GPT-5.6 Preview System Card (system card, 60.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL)
  5. Claude Opus 5 Anthropic0.598as of 2026-07
    length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%).
    System Card: Claude Opus 5 · system card · first-party
    publisher
    Anthropic
    locator
    p. 189, section 8.15.2 HealthBench Professional results; also Table 8.1.A p. 152 ('HealthBench Professional 59.8 ...')
    reported by
    vendor-reported
    published
    2026-07-24
    retrieved
    2026-09-07
    confidence
    verified
    Claude Opus 5 achieved a raw score of 73.4%, which is the highest amongst all Claude models, ahead of Claude Mythos 5 at 70.3%, Claude Opus 4.8 at 60.3%, and Claude Sonnet 5 at 62.4%. After length adjustment, which penalizes verbose model responses, Claude Opus 5 achieved a score of 59.8%.
    also reported in Claude Fable 5.1 and Claude Mythos 5.1 System Card (system card, 59.8%, p. 167, Table 8.1.A, column 'Claude Opus 5'; raw 73.4% also reprinted on p. 199, sec. 8.17.2)
  6. Muse Spark 1.1 Meta0.593as of 2026-07
    length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44)
    Muse Spark 1.1 Evaluation Report · model card · first-party
    publisher
    Meta
    locator
    p. 101, Figure 44 'General capability benchmark results' (image), row HealthBench Professional, column Muse Spark 1.1; protocol p. 104 (printed 103): HealthBench Pro comprises 525 evaluation data points graded by rubrics. We use GPT-5.4 with low reasoning effort as the grader and report the length-normalized rubric score as done in their paper.
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
  7. Claude Sonnet 5 Anthropic0.578as of 2026-06
    length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).
    System Card: Claude Sonnet 5 · system card · first-party
    publisher
    Anthropic
    locator
    p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 5; Figure 8.12.2.A p. 139
    reported by
    vendor-reported
    published
    2026-06-30
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional 57.8 44.2 51.8 -
    also reported in System Card: Claude Opus 5 (system card, 57.8%, p. 189, section 8.15.2, Figure 8.15.2.A)
  8. GPT-5.6 Terra OpenAI0.577as of 2026-06
    length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    also reported in GPT-5.6 Preview System Card (system card, 57.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA)
  9. Claude Opus 4.8 Anthropic0.558as of 2026-05
    length-adjusted, adaptive thinking at max effort, Claude Sonnet 4.6 grader (the grader used by the Opus 4.8 card itself); the other Claude rows on this board use the Claude Opus 4.8 grader, under which Anthropic later prints 57.4 for this model.
    System Card: Claude Opus 4.8 · system card · first-party
    publisher
    Anthropic
    locator
    p. 228, section 8.14.1 HealthBench Professional; Figure 8.14.A p. 229
    reported by
    vendor-reported
    published
    2026-05-28
    retrieved
    2026-09-07
    confidence
    verified
    Claude Opus 4.8 scores 55.8%, a meaningful improvement over Claude Opus 4.7 at 51.9% and Claude Sonnet 4.6 at 41.7%.
    also reported in Claude Fable 5 and Claude Mythos 5 System Card (system card, 56.9, conflicting, p. 252, Table 8.1.A, column Opus 4.8); System Card: Claude Opus 5 (system card, 57.4, conflicting, p. 152, Table 8.1.A, column Opus 4.8)
  10. GPT-5.6 Luna OpenAI0.557as of 2026-06
    length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    also reported in GPT-5.6 Preview System Card (system card, 55.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA)
  11. Muse Spark Meta0.541as of 2026-07
    length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44
    Muse Spark 1.1 Evaluation Report · model card · first-party
    publisher
    Meta
    locator
    p. 101, Figure 44 (image), row HealthBench Professional, column Muse Spark; protocol p. 104 (printed 103)
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
  12. GPT-5.6 Sol (August) OpenAI0.540as of 2026-08
    ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars)
    publisher
    OpenAI
    locator
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional 32.9 (33.8, 2,285) 38.4 (40.7, 2,775) 54.0 (56.6, 2,894) 44.1 (46.8, 2,920)
  13. Claude Opus 4.7 Anthropic0.519as of 2026-05
    length-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 trials; comparison model in the Opus 4.8 card
    System Card: Claude Opus 4.8 · system card · first-party
    publisher
    Anthropic
    locator
    p. 228, section 8.14.1 HealthBench Professional; Figure 8.14.A p. 229
    reported by
    vendor-reported
    published
    2026-05-28
    retrieved
    2026-09-07
    confidence
    verified
    Claude Opus 4.8 scores 55.8%, a meaningful improvement over Claude Opus 4.7 at 51.9% and Claude Sonnet 4.6 at 41.7%.
  14. GPT-5.5 OpenAI0.518as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (57.2 unadjusted, 3818 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    also reported in GPT-5.6 Preview System Card (system card, 51.8, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.5)
  15. GPT-5.4 OpenAI0.481as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.4
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    also reported in GPT-5.6 Preview System Card (system card, 48.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.4)
  16. GPT-5 OpenAI0.462as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    also reported in GPT-5.6 Preview System Card (system card, 46.2, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5)
  17. GPT-5.2 OpenAI0.459as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    also reported in GPT-5.6 Preview System Card (system card, 45.9, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.2)
  18. Claude Sonnet 4.6 Anthropic0.442as of 2026-06
    length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Sonnet 5 card. Other Anthropic prints: 44.4% (Fable card Figure 8.18.2.A), 41.7% (Opus 4.8 card, Sonnet 4.6 grader)
    System Card: Claude Sonnet 5 · system card · first-party
    publisher
    Anthropic
    locator
    p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 4.6
    reported by
    vendor-reported
    published
    2026-06-30
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional 57.8 44.2 51.8 -
    also reported in System Card: Claude Opus 4.8 (system card, 41.7%, conflicting, p. 228, section 8.14.1)
  19. GPT-5.6 Luna (August) OpenAI0.441as of 2026-08
    ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars)
    publisher
    OpenAI
    locator
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional 32.9 (33.8, 2,285) 38.4 (40.7, 2,775) 54.0 (56.6, 2,894) 44.1 (46.8, 2,920)
  20. GPT-5.1 OpenAI0.396as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.1
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    also reported in GPT-5.6 Preview System Card (system card, 39.6, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.1)
  21. GPT-5.5 Instant OpenAI0.384as of 2026-05
    length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
    GPT-5.5 Instant System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT
    reported by
    vendor-reported
    published
    2026-05-05
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Professional 37.6 (40.4, 2,973) 35.7 (38.3, 2,872) 32.9 (33.8, 2,285) 38.4 (40.7, 2,775)
    also reported in GPT-5.6 - August Updates (system card addendum) (system card, 38.4, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant)
  22. MAI-Thinking-1 Microsoft0.350as of 2026-08
    length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision.
    publisher
    Microsoft AI
    locator
    p. 54, Table 12 'Post-trained model evaluation results on various public benchmarks', Health group, column HealthBench Prof.; protocol Appendix K.6 p. 106: HealthBench Professional introduces a length penalty for the primary metric, to correct for a well-observed correlation between lengthy responses and artificially increased LLM-grader scores. For all reported scores, we use the standard GPT-5.4 grader and rubrics provided by OpenAI.
    reported by
    vendor-reported
    published
    2026-08-12
    retrieved
    2026-09-07
    confidence
    verified
    Model AIR-Bench CyberSec Instruct CyberSec Auto Long Fact Truthful QA HealthBench Prof. MedXpert QA MAI-Thinking-1 88 63 63 98 88 35 43 Sonnet 4.6 88 62 56 98 88 38 49

HealthBench Hard

16 rows · 0 to 1

Board source: healthbenchhard.ai · paper: arxiv.org · full board: healthbenchhard.ai

  1. Muse Spark Meta0.428as of 2026-04
    raw score (no length adjustment), GPT-4.1 grader via the OpenAI simple-evals implementation, Muse Spark Thinking; Meta run, launch-post benchmark table.
    publisher
    Meta
    locator
    Launch-post benchmark table image, HEALTH section, row HealthBench Hard, column Muse Spark Thinking; identical table in the Eval Methodology PDF p. 5; protocol p. 2: HealthBench Hard: This is a subset of OpenAI's HealthBench benchmark, containing 1000 prompts. We used the same implementation as in the OpenAI’s official simple-evals repo, with GPT-4.1-genai as the LLM-as-judge model.
    reported by
    vendor-reported
    published
    2026-04-08
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard | Open-Ended Health Queries | 42.8 | 14.8 | 20.6 | 40.1 | 20.3 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT-5.4 Xhigh, Grok 4.2 Reasoning)
    also reported in Muse Spark Eval Methodology (model card, 42.8, p. 5 results table image, row HealthBench Hard; protocol p. 2)
  2. GPT-6 Astra OpenAI0.363as of 2026-09
    length-adjusted, max reasoning effort (37.8 unadjusted, 2,192 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'.
    GPT-6 Astra System Card · system card · first-party
    publisher
    OpenAI
    locator
    p. 19, sec. 6.1 (Table 6 cell, column 'gpt-6 Astra': 36.3 (37.8, 2192))
    reported by
    vendor-reported
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    Astra has a length-adjusted HealthBench Professional score of 63.4 (+2.9 relative to GPT-5.6 Sol), HealthBench score of 58.1 (+1.1), HealthBench Hard score of 36.3 (+3.2), and HealthBench Consensus score of 95.8 (+0.3).
    also reported in GPT-6 Astra System Card - HealthBench (Deployment Safety Hub) (system card, 36.3, mirror, section 6.1, HTML rendering of the same card)
  3. GPT-5 OpenAI0.347as of 2026-06
    length-adjusted, max reasoning effort (41.6 unadjusted, 2,880 mean response chars); GPT-5.6 system card Table 6, column GPT-5. The GPT-5 launch system card printed 46.2% raw for gpt-5-thinking.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
    also reported in GPT-5.6 Preview System Card (system card, 34.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5); GPT-5.5 System Card (system card, 34.7, Section 5 Health, Table 7, column GPT-5); GPT-5 System Card (system card, 46.2, conflicting, p. 18, section 3.10 Health, Figure 6 (HealthBench Hard, raw score %))
  4. GPT-5.2 OpenAI0.343as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (38.9 unadjusted, 2585 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
    also reported in GPT-5.6 Preview System Card (system card, 34.3, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.2)
  5. GPT-5.6 Sol OpenAI0.331as of 2026-06
    length-adjusted, max reasoning effort (31.1 unadjusted, 1,751 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 490.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
    also reported in GPT-5.6 Preview System Card (system card, 33.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL)
  6. GPT-5.6 Terra OpenAI0.327as of 2026-06
    length-adjusted, max reasoning effort (34.3 unadjusted, 2,199 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
    also reported in GPT-5.6 Preview System Card (system card, 32.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA)
  7. GPT-5.6 Luna OpenAI0.320as of 2026-06
    length-adjusted, max reasoning effort (31.4 unadjusted, 1,923 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 491.
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
    also reported in GPT-5.6 Preview System Card (system card, 32.0, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA)
  8. GPT-5.5 OpenAI0.315as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (33.8 unadjusted, 2289 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
    also reported in GPT-5.6 Preview System Card (system card, 31.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.5)
  9. GPT-5.6 Sol (August) OpenAI0.314as of 2026-08
    ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (27.1 unadjusted, 1,450 chars)
    publisher
    OpenAI
    locator
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard 20.2 (17.8, 1,693) 22.9 (21.3, 1,794) 31.4 (27.1, 1,450) 28.7 (24.9, 1,523)
  10. GPT OSS 120B OpenAI0.300as of 2025-08
    raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.
    gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
    publisher
    OpenAI
    locator
    Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-120b high
    reported by
    vendor-reported
    published
    2025-08-05
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard 22.8 26.9 30.0 9.0 12.9 10.8
  11. GPT-5.4 OpenAI0.291as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (30.3 unadjusted, 2161 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.4
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
    also reported in GPT-5.6 Preview System Card (system card, 29.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.4)
  12. GPT-5.6 Luna (August) OpenAI0.287as of 2026-08
    ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (24.9 unadjusted, 1,523 chars)
    publisher
    OpenAI
    locator
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard 20.2 (17.8, 1,693) 22.9 (21.3, 1,794) 31.4 (27.1, 1,450) 28.7 (24.9, 1,523)
  13. GPT-5.3 Chat OpenAI0.259as of 2026-03
    raw score (no length adjustment), GPT-5.3 Instant system card Table 3, column GPT-5.3-INSTANT; OpenAI later cards print 20.2 length-adjusted (17.8 unadjusted) for the re-run model.
    GPT-5.3 Instant System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 4.1 HealthBench, Table 3: HealthBench, row Hard, column GPT-5.3-INSTANT
    reported by
    vendor-reported
    published
    2026-03-02
    retrieved
    2026-09-07
    confidence
    verified
    Hard 26.8% 25.9%
    also reported in GPT-5.5 Instant System Card (system card, 20.2, conflicting, Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.3 INSTANT); GPT-5.6 - August Updates (system card addendum) (system card, 20.2, conflicting, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.3 Instant)
  14. GPT-5.1 OpenAI0.254as of 2026-06
    length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (41.4 unadjusted, 4049 chars)
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.1
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
    also reported in GPT-5.6 Preview System Card (system card, 25.4, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.1)
  15. GPT-5.5 Instant OpenAI0.229as of 2026-05
    length-adjusted (21.3 unadjusted, 1,794 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
    GPT-5.5 Instant System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT
    reported by
    vendor-reported
    published
    2026-05-05
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard 21.6 (23.0, 2,181) 23.3 (23.5, 2,022) 20.2 (17.8, 1,693) 22.9 (21.3, 1,794)
    also reported in GPT-5.6 - August Updates (system card addendum) (system card, 22.9, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant)
  16. GPT OSS 20B OpenAI0.108as of 2025-08
    raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.
    gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
    publisher
    OpenAI
    locator
    Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-20b high
    reported by
    vendor-reported
    published
    2025-08-05
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench Hard 22.8 26.9 30.0 9.0 12.9 10.8

HealthBench

18 rows · 0-100 rubric-point percentage (some sites display 0-1)

Board source: GPT-5.6 System Card · official page: deploymentsafety.openai.com

  1. Claude Opus 5 Anthropic67.1as of 2026-07
    raw/unadjusted, Anthropic system-card protocol (max effort, no tools, five trials); length-adjusted 57.8; mirrored on benchlm.ai
    System Card: Claude Opus 5 · system card · first-party
    publisher
    Anthropic
    locator
    p. 188, section 8.15.1 HealthBench results; Figure 8.15.1.A
    reported by
    vendor-reported
    published
    2026-07-24
    retrieved
    2026-09-07
    confidence
    verified
    Claude Opus 5 achieved a raw score of 67.1%, which is the highest among all Claude models, ahead of Claude Mythos 5 at 62.5%, Claude Opus 4.8 at 58.8%, and Claude Sonnet 5 at 59.2%. After length adjustment, which penalizes verbose model responses, Claude Opus 5 achieved a score of 57.8%.
    also reported in Claude Fable 5.1 and Claude Mythos 5.1 System Card (system card, 67.1%, p. 198, sec. 8.17.1); GPT-5.6 Preview System Card (system card, 67.1)
  2. Baichuan-M3 Baichuan65.1as of 2026-02
    self-run in Baichuan-M3 paper (arXiv 2602.06570)
    publisher
    Baichuan
    locator
    p. 23, section 4.2.1 HealthBench-Main; Table on p. 25 (Model / HealthBench Score)
    reported by
    vendor-reported
    published
    2026-02-06
    retrieved
    2026-09-07
    confidence
    verified
    On the comprehensive HealthBench Total, Baichuan-M3 achieves a score of 65.1, surpassing the runner-up GPT-5.2-High (63.3) by a clear margin.
    also reported in Baichuan-M3 Technical Report (paper, 65.1, p. 23, section 4.2.1 HealthBench-Main (Figure 7); also Table 2, p. 25 (Baichuan-M3-235B, HealthBench Score 65.1))
  3. GPT-5.2-High OpenAI63.3as of 2026-02
    raw score as run by Baichuan in the M3 technical report, not an OpenAI-reported number; OpenAI own GPT-5.2 figure is 56.8 length-adjusted (60.7 unadjusted) in the GPT-5.6 system card.
    publisher
    Baichuan
    locator
    p. 25, HealthBench-Hallu table (Model / HealthBench Score column); also p. 23 prose
    reported by
    independent run
    published
    2026-02-06
    retrieved
    2026-09-07
    confidence
    verified
    GPT-5.2-High 63.3 2.37% 2.78%
    also reported in Baichuan-M3 Technical Report (paper, 63.3, p. 23, section 4.2.1 HealthBench-Main (Figure 7); also Table 2, p. 25 (GPT-5.2-High, HealthBench Score 63.3)); GPT-5.6 System Card (system card, 56.8, conflicting, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2)
  4. Claude Fable 5 Anthropic62.7as of 2026-06
    length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (40k token budget); printed in the Mythos 5 column of Table 8.1.A (Fable 5 column '-'); raw 62.5% per Opus 5 card
    publisher
    Anthropic
    locator
    p. 252, Table 8.1.A, row HealthBench, column Mythos 5; Figure 8.18.1.A p. 297
    reported by
    vendor-reported
    published
    2026-06-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench 62.7 - 61.1 59.3 56.5 -
    also reported in System Card: Claude Opus 5 (system card, 62.7%, p. 188, section 8.15.1, Figure 8.15.1.A); Claude Fable 5.1 and Claude Mythos 5.1 System Card (system card, 61.2% (raw), conflicting, p. 198, sec. 8.17.1)
  5. Claude Opus 4.8 Anthropic59.3as of 2026-06
    length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Fable/Mythos 5 card; raw 58.8% per Opus 5 card; not in the Opus 4.8 card itself
    publisher
    Anthropic
    locator
    p. 252, Table 8.1.A, row HealthBench, column Opus 4.8
    reported by
    vendor-reported
    published
    2026-06-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench 62.7 - 61.1 59.3 56.5 -
    also reported in System Card: Claude Opus 5 (system card, 59.3%, p. 188, section 8.15.1, Figure 8.15.1.A)
  6. Claude Sonnet 5 Anthropic58.7%as of 2026-06
    length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; figure-only in the Sonnet 5 card; raw 59.2% per Opus 5 card
    System Card: Claude Sonnet 5 · system card · first-party
    publisher
    Anthropic
    locator
    p. 138, section 8.12.1 HealthBench results, Figure 8.12.1.A bar label (no prose or table number)
    reported by
    vendor-reported
    published
    2026-06-30
    retrieved
    2026-09-07
    confidence
    verified
    [Figure 8.12.1.A] HealthBench length-adjusted scores. All Claude models used adaptive thinking at max effort. (bar label: Claude Sonnet 5 58.7%)
    also reported in System Card: Claude Opus 5 (system card, 58.7%, p. 188, section 8.15.1, Figure 8.15.1.A)
  7. GPT-5.3 Chat OpenAI54.1%as of 2026-03
    raw score (no length adjustment), GPT-5.3 Instant system card Table 3 column GPT-5.3-INSTANT; later OpenAI cards print 49.6 length-adjusted (47.9 unadjusted)
    GPT-5.3 Instant System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 4.1 HealthBench, Table 3: HealthBench, row HealthBench, column GPT-5.3-INSTANT
    reported by
    vendor-reported
    published
    2026-03-02
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench 55.4% 54.1%
    also reported in GPT-5.5 Instant System Card (system card, 49.6, conflicting, Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.3 INSTANT)
  8. Claude Fable 5.1 Anthropic60%as of 2026-09
    length-adjusted (method published in OpenAI's GPT-5.5 System Card); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 66.7%).
    publisher
    Anthropic
    locator
    p. 198, sec. 8.17.1 / Figure 8.17.1.A
    reported by
    vendor-reported
    published
    2026-09-01
    retrieved
    2026-09-07
    confidence
    verified
    On HealthBench, Claude Fable 5.1 achieved a raw score of 66.7%, ahead of Claude Fable 5 at 61.2% and Claude Sonnet 5 at 59.2%, and behind Claude Opus 5 at 67.1%. After length adjustment, which penalizes verbose model responses, Fable 5.1 achieved a score of 60%.
  9. GPT-6 Astra OpenAI58.1as of 2026-09
    length-adjusted, max reasoning effort (59.7 unadjusted, 2,258 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'.
    GPT-6 Astra System Card · system card · first-party
    publisher
    OpenAI
    locator
    p. 19, sec. 6.1 (Table 6 cell, column 'gpt-6 Astra': 58.1 (59.7, 2258))
    reported by
    vendor-reported
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    Astra has a length-adjusted HealthBench Professional score of 63.4 (+2.9 relative to GPT-5.6 Sol), HealthBench score of 58.1 (+1.1), HealthBench Hard score of 36.3 (+3.2), and HealthBench Consensus score of 95.8 (+0.3).
    also reported in GPT-6 Astra System Card - HealthBench (Deployment Safety Hub) (system card, 58.1, mirror, section 6.1, HTML rendering of the same card)
  10. GPT OSS 120B OpenAI57.6as of 2025-08
    reasoning level high, raw score (%), gpt-oss model card Table 3 (low 53.0, medium 55.9)
    gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
    publisher
    OpenAI
    locator
    Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench, column gpt-oss-120b high
    reported by
    vendor-reported
    published
    2025-08-05
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench 53.0 55.9 57.6 40.4 41.8 42.5
  11. GPT-5.6 Sol OpenAI57.0as of 2026-06
    length-adjusted, max reasoning effort (55.6 unadjusted), GPT-5.6 system card 2026-07-09
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench length-adjusted 57.7 (63.1, 2904) 50.9 (64.2, 4222) 56.8 (60.7, 2645) 54.0 (55.7, 2275) 56.5 (58.4, 2313) 57.0 (55.6, 1764) 57.0 (58.7, 2285) 55.8 (55.4, 1930)
    also reported in GPT-5.6 Preview System Card (system card, 57.0, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL)
  12. GPT-5.6 Terra OpenAI57.0as of 2026-06
    length-adjusted (58.7 unadjusted), max reasoning effort
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench length-adjusted 57.7 (63.1, 2904) 50.9 (64.2, 4222) 56.8 (60.7, 2645) 54.0 (55.7, 2275) 56.5 (58.4, 2313) 57.0 (55.6, 1764) 57.0 (58.7, 2285) 55.8 (55.4, 1930)
    also reported in GPT-5.6 Preview System Card (system card, 57.0, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA)
  13. GPT-5.5 OpenAI56.5as of 2026-04
    length-adjusted (58.4 unadjusted), comparison row in GPT-5.6 system card
    GPT-5.5 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5 Health, Table 7 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
    reported by
    vendor-reported
    published
    2026-04-23
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench length-adjusted 57.7 (63.1, 2904) 50.9 (64.2, 4222) 56.8 (60.7, 2645) 54.0 (55.7, 2275) 56.5 (58.4, 2313)
    also reported in GPT-5.6 Preview System Card (system card, 56.5); GPT-5.6 System Card (system card, 56.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5)
  14. GPT-5.6 Luna OpenAI55.8as of 2026-06
    length-adjusted (55.4 unadjusted), max reasoning effort
    GPT-5.6 System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench length-adjusted 57.7 (63.1, 2904) 50.9 (64.2, 4222) 56.8 (60.7, 2645) 54.0 (55.7, 2275) 56.5 (58.4, 2313) 57.0 (55.6, 1764) 57.0 (58.7, 2285) 55.8 (55.4, 1930)
    also reported in GPT-5.6 Preview System Card (system card, 55.8, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA)
  15. GPT-5.6 Sol (August) OpenAI55.0as of 2026-08
    ChatGPT production/Instant deployment setting, length-adjusted (52.1 unadjusted), GPT-5.6 August Updates PDF 2026-08-06
    publisher
    OpenAI
    locator
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench 49.6 (47.9, 1,724) 51.4 (50.9, 1,922) 55.0 (52.1, 1,514) 53.3 (50.7, 1,567)
    also reported in GPT-5.6 Preview System Card (system card, 55.0)
  16. GPT-5.6 Luna (August) OpenAI53.3as of 2026-08
    ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (50.7 unadjusted, 1,567 chars)
    publisher
    OpenAI
    locator
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench 49.6 (47.9, 1,724) 51.4 (50.9, 1,922) 55.0 (52.1, 1,514) 53.3 (50.7, 1,567)
  17. GPT-5.5 Instant OpenAI51.4as of 2026-05
    length-adjusted, GPT-5.5 Instant system card Table 5 column GPT-5.5 INSTANT (50.9 unadjusted, 1,922 chars); same number in GPT-5.6 August Updates p. 11
    GPT-5.5 Instant System Card · system card · first-party
    publisher
    OpenAI
    locator
    Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT
    reported by
    vendor-reported
    published
    2026-05-05
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench 49.6 (50.8, 2,208) 50.6 (51.5, 2,145) 49.6 (47.9, 1,724) 51.4 (50.9, 1,922)
    also reported in GPT-5.6 - August Updates (system card addendum) (system card, 51.4, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant)
  18. GPT OSS 20B OpenAI42.5as of 2025-08
    reasoning level high, raw score (%), gpt-oss model card Table 3 (low 40.4, medium 41.8)
    gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
    publisher
    OpenAI
    locator
    Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench, column gpt-oss-20b high
    reported by
    vendor-reported
    published
    2025-08-05
    retrieved
    2026-09-07
    confidence
    verified
    HealthBench 53.0 55.9 57.6 40.4 41.8 42.5

Health Optimization Bench

16 rows · 0-100 rubric credit

Board source: healthoptimizationbench.com · full board: healthoptimizationbench.com

  1. Claude Fable 5.1 Anthropic84.9as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 84.9, 95% CI 80.1 to 89.3, n 89 (values as published in the site snapshot of 2026-09-04)
  2. Claude Fable 5 Anthropic83.8as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 83.8, 95% CI 79.1 to 88.1, n 89 (values as published in the site snapshot of 2026-09-04)
  3. Grok 4.6 xAI81.3as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 81.3, 95% CI 77 to 85.3, n 89 (values as published in the site snapshot of 2026-09-04)
  4. Claude Opus 5 Anthropic78.3as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 78.3, 95% CI 72.9 to 83.4, n 89 (values as published in the site snapshot of 2026-09-04)
  5. GPT-5.6 Sol (max) OpenAI77.9as of 2026-09
    max reasoning effort
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 77.9, 95% CI 73.3 to 82.2, n 89 (values as published in the site snapshot of 2026-09-04)
  6. Kimi K3 Moonshot AI77.8as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 77.8, 95% CI 73 to 82.4, n 89 (values as published in the site snapshot of 2026-09-04)
  7. GPT-5.6 Sol (high) OpenAI77.7as of 2026-09
    high reasoning effort
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 77.7, 95% CI 73.1 to 82, n 89 (values as published in the site snapshot of 2026-09-04)
  8. Muse Spark Meta68.4as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 68.4, 95% CI 62.8 to 73.9, n 89 (values as published in the site snapshot of 2026-09-04)
  9. Gemini 3.6 Google57.9as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 57.9, 95% CI 51.2 to 64.3, n 89 (values as published in the site snapshot of 2026-09-04)
  10. Inkling Thinking Machines53.9as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 53.9, 95% CI 47.7 to 60.1, n 89 (values as published in the site snapshot of 2026-09-04)
  11. Claude Sonnet 5 Anthropic50.1as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 50.1, 95% CI 44.3 to 56.1, n 89 (values as published in the site snapshot of 2026-09-04)
  12. MiniMax M3 MiniMax35.5as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 35.5, 95% CI 29.6 to 41.6, n 89 (values as published in the site snapshot of 2026-09-04)
  13. MAI Thinking Microsoft AI33.0as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 33.0, 95% CI 26.8 to 39.3, n 89 (values as published in the site snapshot of 2026-09-04)
  14. GLM 5.2 Zhipu29.2as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 29.2, 95% CI 23.9 to 34.9, n 89 (values as published in the site snapshot of 2026-09-04)
  15. Mistral Medium 3.5 Mistral14.4as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 14.4, 95% CI 10.6 to 18.5, n 89 (values as published in the site snapshot of 2026-09-04)
  16. Nemotron 3.5 Lightning NVIDIA7.8as of 2026-09
    publisher
    Arcophos
    locator
    tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
    reported by
    Arcophos run
    retrieved
    not recorded
    confidence
    verified
    score 7.8, 95% CI 4.7 to 11.3, n 89 (values as published in the site snapshot of 2026-09-04)

MAST (Medical AI Superintelligence Test)

8 rows · percentage composite

Board source: MAST: Medical AI Superintelligence Test leaderboard (General board)

  1. GPT-5.6 Sol OpenAI60.2%as of 2026-08
    MAST in preview; 'exact scores may change'
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    1 GPT-5.6 Sol OpenAI 60.2%
  2. Kimi K3 Moonshot AI60.1%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    2 Kimi K3 Moonshot AI 60.1%
  3. Gemini 3.6 Flash Google59.3%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    3 Gemini 3.6 Flash Google 59.3%
  4. Gemini 3.1 Pro Google58.9%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    4 Gemini 3.1 Pro Google 58.9%
  5. Qwen3.5 397B A17B Alibaba57.9%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    5 Qwen3.5 397B A17B Alibaba 57.9%
  6. Claude Opus 5 Anthropic57.1%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    6 Claude Opus 5 Anthropic 57.1%
  7. Claude Sonnet 5 Anthropic56.6%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    7 Claude Sonnet 5 Anthropic 56.6%
  8. Grok 4.3 xAI53.7%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    8 Grok 4.3 xAI 53.7%

MedHELM

10 rows · mean win rate 0-1

Board source: MedHELM leaderboard (medhelm.org), v5.0.0

  1. Gemini 3.1 Pro (Preview) Google0.652as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-07
    confidence
    verified
    1 Gemini 3.1 Pro (Preview) Google 0.652
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.6520833333333333, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  2. Gemini 3.5 Flash Google0.642as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-07
    confidence
    verified
    2 Gemini-3.5-flash Google 0.642
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.6416666666666667, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  3. Muse Spark (2026-04-08) Meta0.621as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-07
    confidence
    verified
    3 Muse Spark (2026-04-08) Meta 0.621
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.6208333333333333, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  4. GPT-5.4 mini OpenAI0.552as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-07
    confidence
    verified
    4 GPT-5.4 mini OpenAI 0.552
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.5520833333333334, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  5. GPT-5.4 (2026-03-05) OpenAI0.538as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-07
    confidence
    verified
    5 GPT-5.4 (2026-03-05) OpenAI 0.538
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.5375, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  6. Gemini 2.5 Pro Google0.529as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-07
    confidence
    verified
    6 Gemini 2.5 Pro (05-06 preview) Google 0.529
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.5291666666666667, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  7. DeepSeek R1 DeepSeek0.485as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-07
    confidence
    verified
    7 DeepSeek R1 DeepSeek 0.485
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.48541666666666666, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  8. Claude 4.6 Opus Anthropic0.456as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-07
    confidence
    verified
    8 Claude 4.6 Opus Anthropic 0.456
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.45625, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  9. Claude 3.7 Sonnet Anthropic0.45as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-07
    confidence
    verified
    9 Claude 3.7 Sonnet (20250219) Anthropic 0.45
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.45, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
  10. Gemini 2.0 Flash Google0.342as of 2026-05
    MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
    publisher
    Stanford CRFM (MedHELM)
    locator
    medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    reported by
    official leaderboard
    published
    2026-05-14
    retrieved
    2026-09-07
    confidence
    verified
    10 Gemini 2.0 Flash Google 0.342
    also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.3416666666666667, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')

First, Do NOHARM (v2)

12 rows · percentage safety score

Board source: MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · paper: arxiv.org

  1. Muse Spark 1.1 Meta79.7%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    5Muse Spark 1.1Meta 79.7%
  2. Claude Opus 5 Anthropic74.6%as of 2026-08
    v2 run on ARISE; 19 models on the board
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    6Claude Opus 5Anthropic 74.6%
  3. Kimi K3 Moonshot AI74.0%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    7Kimi K3OSSMoonshot AI 74.0%
  4. GPT-5.6 Sol OpenAI70.1%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    8GPT-5.6 SolOpenAI 70.1%
  5. GPT-5.5 OpenAI70.0%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    9GPT-5.5OpenAI 70.0%
  6. GPT-5 OpenAI68.6%as of 2026-08
    from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI); not in the Latest Flagships ranking
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, Model Leaderboard (Top 10 shown), row 10, SAFETY column
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    10 GPT-5 72.6%±2.3 68.6%±4.5 80.3%±4.4 30.1%±9.0 44.2%±2.5 45.4%±1.3 73.1%±3.9 73.7%±2.7 46.6%±0.0
  7. Claude Fable 5 Anthropic65.0%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    10Claude Fable 5Anthropic 65.0%
  8. Gemini 3.1 Pro Google62.6%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    11Gemini 3.1 ProGoogle 62.6%
  9. Gemini 2.5 Pro Google61.9%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    12Gemini 2.5 ProGoogle 61.9%
  10. Qwen3.5 397B A17B Alibaba61.1%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    13Qwen3.5 397B A17BOSSAlibaba 61.1%
  11. Kimi K2.6 Moonshot AI59.1%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    14Kimi K2.6OSSMoonshot AI 59.1%
  12. DeepSeek R1 DeepSeek55.8%as of 2026-08
    publisher
    ARISE AI Research Network
    locator
    arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    reported by
    official leaderboard
    published
    2026-08-15
    retrieved
    2026-09-07
    confidence
    verified
    16DeepSeek R1OSSDeepSeek 55.8%

HealthAgentBench

11 rows · mean task success rate

Board source: HealthAgentBench leaderboard · paper: arxiv.org

  1. Claude Code (Opus 5) Anthropic55%as of 2026-07
    $3.3/task; harness+model evaluated jointly
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 1, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-07
    confidence
    verified
    1 Claude Code (Opus 5) 55% $3.3
    also reported in HealthAgentBench detailed results (official leaderboard, 55%, Detailed results table, rank 1)
  2. Codex (GPT-5.6-sol) OpenAI45%as of 2026-07
    $5.2/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 2, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-07
    confidence
    verified
    2 Codex (GPT-5.6-sol) 45% $5.2
    also reported in HealthAgentBench detailed results (official leaderboard, 45%, Detailed results table, rank 2)
  3. Codex (GPT 5.5) OpenAI42%as of 2026-07
    $2.8/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 3, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-07
    confidence
    verified
    3 Codex (GPT 5.5) 42% $2.8
    also reported in HealthAgentBench detailed results (official leaderboard, 42%, Detailed results table, rank 3); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 42%, p. 11, Figure 4 (pooled task success rate, ten agents))
  4. Copilot (Opus 4.8) Microsoft/Anthropic36%as of 2026-07
    $3.1/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 4, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-07
    confidence
    verified
    4 Copilot (Opus 4.8) 36% $3.1
    also reported in HealthAgentBench detailed results (official leaderboard, 36%, Detailed results table, rank 4); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 36%, p. 11, Figure 4 (pooled task success rate, ten agents))
  5. Copilot (GPT 5.5) Microsoft/OpenAI35%as of 2026-07
    $2.6/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 5, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-07
    confidence
    verified
    5 Copilot (GPT 5.5) 35% $2.6
    also reported in HealthAgentBench detailed results (official leaderboard, 35%, Detailed results table, rank 5); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 35%, p. 11, Figure 4 (pooled task success rate, ten agents))
  6. Claude Code (Opus 4.8) Anthropic32%as of 2026-07
    $4.0/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 6, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-07
    confidence
    verified
    6 Claude Code (Opus 4.8) 32% $4.0
    also reported in HealthAgentBench detailed results (official leaderboard, 32%, Detailed results table, rank 6); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 32%, p. 11, Figure 4 (pooled task success rate, ten agents))
  7. Codex (GPT 5.4) OpenAI28%as of 2026-07
    $1.3/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 7, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-07
    confidence
    verified
    7 Codex (GPT 5.4) 28% $1.3
    also reported in HealthAgentBench detailed results (official leaderboard, 28%, Detailed results table, rank 7); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 28%, p. 11, Figure 4 (pooled task success rate, ten agents))
  8. Claude Code (Opus 4.7) Anthropic27%as of 2026-07
    $4.8/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 8, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-07
    confidence
    verified
    8 Claude Code (Opus 4.7) 27% $4.8
    also reported in HealthAgentBench detailed results (official leaderboard, 27%, Detailed results table, rank 8); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 27%, p. 11, Figure 4 (pooled task success rate, ten agents))
  9. Claude Code (Opus 4.6) Anthropic19%as of 2026-07
    $4.1/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 10, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-07
    confidence
    verified
    10 Claude Code (Opus 4.6) 19% $4.1
    also reported in HealthAgentBench detailed results (official leaderboard, 19%, Detailed results table, rank 10); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 19%, p. 11, Figure 4 (pooled task success rate, ten agents))
  10. Claude Code (Sonnet 4.6) Anthropic17%as of 2026-07
    $2.9/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 11, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-07
    confidence
    verified
    11 Claude Code (Sonnet 4.6) 17% $2.9
    also reported in HealthAgentBench detailed results (official leaderboard, 17%, Detailed results table, rank 11); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 17%, p. 11, Figure 4 (pooled task success rate, ten agents))
  11. Codex (GPT 5.4 Mini) OpenAI16%as of 2026-07
    $0.6/task
    HealthAgentBench leaderboard · official leaderboard · first-party
    publisher
    Microsoft Research (HealthAgentBench)
    locator
    Leaderboard table (homepage), rank 12, Success Rate column
    reported by
    official leaderboard
    published
    2026-07-27
    retrieved
    2026-09-07
    confidence
    verified
    12 Codex (GPT 5.4 Mini) 16% $0.6
    also reported in HealthAgentBench detailed results (official leaderboard, 16%, Detailed results table, rank 12); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 16%, p. 11, Figure 4 (pooled task success rate, ten agents))

CHI-Bench

12 rows · pass@1 with binary 0/1 reward

Board source: CHI-Bench leaderboard (actAVA)

  1. erius + claude-opus-5 Humana (harness) / Anthropic (model)54.7%as of 2026-08
    community-submitted harness config validated by automated workspace judge
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    Leaderboard table, All Domains, rank 01, Accuracy column
    reported by
    independent run
    published
    2026-08-12
    retrieved
    2026-09-07
    confidence
    verified
    01 | erius | submitted by Michael Johnson (MJ) | claude-opus-5 | Proprietary | 54.7% | 72.0% | 36.0% | 56.0% | 2026-07-26
  2. claude-code + claude-opus-5 Anthropic37.3%as of 2026-08
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    Leaderboard table, All Domains, rank 03, Accuracy column
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-07
    confidence
    verified
    03 | claude-code | claude-opus-5 | Proprietary | 37.3% | 20.0% | 32.0% | 60.0% | 2026-07-24
  3. erius + claude-opus-4-8 Humana / Anthropic37.3%as of 2026-08
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    Leaderboard table, All Domains, rank 02, Accuracy column
    reported by
    independent run
    published
    2026-08-12
    retrieved
    2026-09-07
    confidence
    verified
    02 | erius | submitted by Michael Johnson (MJ) | claude-opus-4-8 | Proprietary | 37.3% | 40.0% | 16.0% | 56.0% | 2026-06-05
  4. claude-code + claude-opus-4-8 Anthropic33.3%as of 2026-08
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    Leaderboard table, All Domains, rank 04, Accuracy column
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-07
    confidence
    verified
    04 | claude-code | claude-opus-4-8 | Proprietary | 33.3% | 32.0% | 28.0% | 40.0% | 2026-05-28
  5. claude-code + claude-opus-4-6 Anthropic28.0%as of 2026-08
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    Leaderboard table, All Domains, rank 05, Accuracy column
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-07
    confidence
    verified
    05 | claude-code | claude-opus-4-6 | Proprietary | 28.0% | 20.0% | 36.0% | 28.0% | 2026-05-01
  6. claude-code + claude-sonnet-4-6 Anthropic26.2%as of 2026-08
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    Leaderboard table, All Domains, rank 06, Accuracy column
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-07
    confidence
    verified
    06 | claude-code | claude-sonnet-4-6 | Proprietary | 26.2% | 24.0% | 34.7% | 20.0% | 2026-05-01
  7. codex + gpt-5.6-sol OpenAI25.3%as of 2026-08
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    Leaderboard table, All Domains, rank 07, Accuracy column
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-07
    confidence
    verified
    07 | codex | gpt-5.6-sol | Proprietary | 25.3% | 36.0% | 28.0% | 12.0% | 2026-07-24
  8. openai-agents + kimi-k3 Moonshot AI25.3%as of 2026-08
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    Leaderboard table, All Domains, rank 08, Accuracy column
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-07
    confidence
    verified
    08 | openai-agents | kimi-k3 | Open-source | 25.3% | 28.0% | 32.0% | 16.0% | 2026-07-24
  9. claude-code + claude-opus-4-7 Anthropic24.4%as of 2026-08
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    Leaderboard table, All Domains, rank 09, Accuracy column
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-07
    confidence
    verified
    09 | claude-code | claude-opus-4-7 | Proprietary | 24.4% | 24.0% | 17.3% | 32.0% | 2026-05-01
  10. claude-code + claude-fable-5 Anthropic24.0%as of 2026-08
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    Leaderboard table, All Domains, rank 10, Accuracy column
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-07
    confidence
    verified
    10 | claude-code | claude-fable-5 | Proprietary | 24.0% | 24.0% | 24.0% | 24.0% | 2026-07-22
  11. codex + gpt-5.5 OpenAI20.9%as of 2026-08
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    Leaderboard table, All Domains, rank 12, Accuracy column
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-07
    confidence
    verified
    12 | codex | gpt-5.5 | Proprietary | 20.9% | 29.3% | 32.0% | 1.3% | 2026-05-01
  12. claude-code + claude-sonnet-5 Anthropic20.0%as of 2026-08
    CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
    publisher
    actAVA
    locator
    Leaderboard table, All Domains, rank 13, Accuracy column
    reported by
    official leaderboard
    published
    2026-08-12
    retrieved
    2026-09-07
    confidence
    verified
    13 | claude-code | claude-sonnet-5 | Proprietary | 20.0% | 24.0% | 24.0% | 12.0% | 2026-07-06

MedCode (Vals AI)

11 rows · percentage accuracy 0-100

Board source: Vals AI MedCode leaderboard

  1. Claude Opus 5 Anthropic63.57%as of 2026-09
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (Overall), rank 1, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    1 Claude Opus 5 63.57% ±1.99 $5/$25 24.50s
  2. Gemini 3.1 Pro Preview (02/26) Google59.06%as of 2026-09
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (Overall), rank 2, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    2 Gemini 3.1 Pro Preview (02/26) 59.06% ±2.00 $2/$12 38.52s
  3. Claude Fable 5 Anthropic56.07%as of 2026-09
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (Overall), rank 3, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    3 Claude Fable 5 56.07% ±2.20 $10/$50 91.44s
  4. Gemini 3 Flash (12/25) Google55.92%as of 2026-09
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (Overall), rank 4, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    4 Gemini 3 Flash (12/25) 55.92% ±2.11 $0.5/$3 44.15s
  5. Gemini 3.5 Flash Google55.83%as of 2026-09
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (Overall), rank 5, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    5 Gemini 3.5 Flash 55.83% ±2.11 $1.5/$9 25.29s
  6. Claude Opus 4.7 Anthropic54.86%as of 2026-09
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (Overall), rank 6, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    6 Claude Opus 4.7 54.86% ±2.21 $5/$25 54.25s
  7. Claude Fable 5.1 Anthropic53.51%as of 2026-09
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (Overall), rank 7, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    7 Claude Fable 5.1 53.51% ±2.17 $10/$50 3m35s
  8. Claude Opus 4.8 Anthropic53.22%as of 2026-09
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (Overall), rank 9, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    9 Claude Opus 4.8 53.22% ±2.17 $5/$25 105.92s
  9. Gemini 3.6 Flash Google53.15%as of 2026-09
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (Overall), rank 10, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    10 Gemini 3.6 Flash 53.15% ±2.16 $1.5/$7.5 16.68s
  10. GPT 5.1 OpenAI52.73%as of 2026-09
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (Overall), rank 11, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    11 GPT 5.1 52.73% ±2.15 $1.25/$10 54.55s
  11. GPT-6 Astra OpenAI48.49%as of 2026-09
    Vals AI MedCode leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (Overall), rank 22 of 90, Accuracy column; board 'Updated 9/3/2026'
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    "openai/gpt-6-astra":[0,{"accuracy":[0,48.486],"latency":[0,102.865],"stderr":[0,2.131],"cost_per_test":[0,0.451358],"temperature":[0,null],"top_p":[0,null],"max_output_tokens":[0,128000]

MedScribe (Vals AI)

12 rows · percentage accuracy 0-100

Board source: Vals AI MedScribe leaderboard

  1. Claude Fable 5.1 Anthropic91.29%as of 2026-09
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (View: All Models, Task: Overall), rank 1, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    1 Claude Fable 5.1 91.29% ±1.95 $10/$50 3m08s
  2. Claude Opus 5 Anthropic90.98%as of 2026-09
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (View: All Models, Task: Overall), rank 2, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    2 Claude Opus 5 90.98% ±1.92 $5/$25 76.56s
  3. Muse Spark 1.2 Meta90.06%as of 2026-09
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (View: All Models, Task: Overall), rank 3, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    3 Muse Spark 1.2 90.06% ±1.96 $1.25/$4.25 61.46s
  4. Muse Spark 1.1 Meta88.89%as of 2026-09
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (View: All Models, Task: Overall), rank 5, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    5 Muse Spark 1.1 88.89% ±1.95 $1.25/$4.25 63.34s
  5. Claude Fable 5 Anthropic88.52%as of 2026-09
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (View: All Models, Task: Overall), rank 7, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    7 Claude Fable 5 88.52% ±1.95 $10/$50 119.47s
  6. GPT 5.1 OpenAI88.09%as of 2026-09
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (View: All Models, Task: Overall), rank 8, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    8 GPT 5.1 88.09% ±1.94 $1.25/$10 77.98s
  7. Kimi K3 Moonshot AI87.96%as of 2026-09
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (View: All Models, Task: Overall), rank 9, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    9 Kimi K3 87.96% ±1.89 $3/$15 2m17s
  8. GPT-6 Astra OpenAI87.91%as of 2026-09
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (View: All Models, Task: Overall), rank 10 of 92, Accuracy column; board 'Updated 9/3/2026'
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    "openai/gpt-6-astra":[0,{"accuracy":[0,87.908],"latency":[0,153.728],"stderr":[0,1.938],"cost_per_test":[0,0.581991],"temperature":[0,null],"top_p":[0,null],"max_output_tokens":[0,128000]
  9. MiniMax-M3 MiniMax87.25%as of 2026-09
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (View: All Models, Task: Overall), rank 11, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    11 MiniMax-M3 87.25% ±1.96 $0.6/$2.4 2m04s
  10. GPT 5.5 OpenAI86.87%as of 2026-09
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (View: All Models, Task: Overall), rank 13, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    13 GPT 5.5 86.87% ±1.93 $5/$30 2m13s
  11. Claude Opus 4.6 (Nonthinking) Anthropic86.74%as of 2026-09
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (View: All Models, Task: Overall), rank 14, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    14 Claude Opus 4.6 (Nonthinking) 86.74% ±1.94 $5/$25 54.32s
  12. Grok 4.6 xAI86.53%as of 2026-09
    Vals AI MedScribe leaderboard · official leaderboard · first-party
    publisher
    Vals AI
    locator
    Leaderboard table (View: All Models, Task: Overall), rank 15, Accuracy column
    reported by
    official leaderboard
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    15 Grok 4.6 86.53% ±1.96 $2/$6 70.94s

MedXpertQA (MM)

15 rows · percentage accuracy 0-100

Board source: Introducing Muse Spark: Scaling Towards Personal Superintelligence · official page: benchlm.ai

  1. GPT-5.6 Sol OpenAI81.5
    Qwen-run comparison in the Qwen3.8-Max launch post
    Qwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
    publisher
    Alibaba
    locator
    Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
    reported by
    independent run
    published
    2026-08-02
    retrieved
    2026-09-07
    confidence
    verified
    | MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
  2. Gemini 3.1 Pro Google81.3%as of 2026-04
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    publisher
    Meta
    locator
    Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Gemini 3.1 Pro High; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    reported by
    independent run
    published
    2026-04-08
    retrieved
    2026-09-07
    confidence
    verified
    MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
    also reported in Muse Spark Eval Methodology (model card, 81.3, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column Gemini 3.1 Pro High); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 81.3%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 1; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  3. Qwen3.8 Max Alibaba80.4%as of 2026-08
    Alibaba's own Qwen3.8 launch table; protocol not stated
    Qwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
    publisher
    Alibaba
    locator
    Multimodal Benchmarks table, Multimodal Reasoning section, row MedXpertQA-MM, column Qwen3.8-Max; header row: | | Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus | Qwen3.8-Max |
    reported by
    vendor-reported
    published
    2026-08-02
    retrieved
    2026-09-07
    confidence
    verified
    | MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
    also reported in MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 80.4%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 2; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  4. Claude Fable 5 Anthropic80.0
    Qwen-run comparison in the Qwen3.8-Max launch post
    Qwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
    publisher
    Alibaba
    locator
    Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
    reported by
    independent run
    published
    2026-08-02
    retrieved
    2026-09-07
    confidence
    verified
    | MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
  5. Muse Spark Meta78.4%as of 2026-04
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    publisher
    Meta
    locator
    Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Muse Spark Thinking; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    reported by
    vendor-reported
    published
    2026-04-08
    retrieved
    2026-09-07
    confidence
    verified
    MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
    also reported in Muse Spark Eval Methodology (model card, 78.4, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column Muse Spark Thinking); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 78.4%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 3; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  6. GPT-5.4 OpenAI77.1%as of 2026-04
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    publisher
    Meta
    locator
    Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column GPT 5.4 Xhigh; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    reported by
    independent run
    published
    2026-04-08
    retrieved
    2026-09-07
    confidence
    verified
    MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
    also reported in Muse Spark Eval Methodology (model card, 77.1, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column GPT 5.4 Xhigh); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 77.1%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 4; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  7. GPT-5.2 OpenAI73.3
    Qwen-run comparison in the Qwen3.5-397B-A17B model card
    Qwen/Qwen3.5-397B-A17B model card · model card · first-party
    publisher
    Hugging Face
    locator
    Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
    reported by
    independent run
    published
    2026-02-16
    retrieved
    2026-09-07
    confidence
    verified
    MedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
  8. Claude Opus 4.8 Anthropic71.7
    Qwen-run comparison in the Qwen3.8-Max launch post
    Qwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
    publisher
    Alibaba
    locator
    Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
    reported by
    independent run
    published
    2026-08-02
    retrieved
    2026-09-07
    confidence
    verified
    | MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
  9. Qwen3.7 Plus Alibaba71.0%as of 2026-05
    Alibaba's own Qwen3.7 Plus launch table; protocol not stated
    Qwen3.7-Plus: Multimodal Agent Intelligence · launch post · first-party
    publisher
    Alibaba
    locator
    Multimodal Benchmarks table, Multimodal Reasoning section, row MedXpertQA-MM, column Qwen3.7-Plus; header row: | | GPT-5.4 (xhigh) | Opus-4.6 Max | Gemini-3.1 Pro | Qwen3.6-Plus | Qwen3.7-Plus |
    reported by
    vendor-reported
    published
    2026-05-31
    retrieved
    2026-09-07
    confidence
    verified
    | MedXpertQA-MM | 77.3 | 64.4 | 80.7 | 68.7 | 71.0 |
    also reported in Qwen3.8-Max: A New Bar for Coding and Cowork (launch post, 71.0, Qwen3.8 launch post, Multimodal Benchmarks table, row MedXpertQA-MM, column Qwen3.7-Plus (same value carried forward)); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 71.0%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 5; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  10. Qwen3.5 397B A17B Alibaba70.0
    self-reported in the Qwen3.5-397B-A17B model card
    Qwen/Qwen3.5-397B-A17B model card · model card · first-party
    publisher
    Hugging Face
    locator
    Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
    reported by
    vendor-reported
    published
    2026-02-16
    retrieved
    2026-09-07
    confidence
    verified
    MedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
  11. Qwen3.6 Plus Alibaba68.7
    Qwen-run comparison in the Qwen3.7-Plus launch post
    Qwen3.7-Plus: Multimodal Agent Intelligence · launch post · first-party
    publisher
    Alibaba
    locator
    Multimodal Benchmarks table, row MedXpertQA-MM; columns GPT-5.4 (xhigh) / Opus-4.6 Max / Gemini-3.1 Pro / Qwen3.6-Plus / Qwen3.7-Plus
    reported by
    vendor-reported
    published
    2026-05-31
    retrieved
    2026-09-07
    confidence
    verified
    | MedXpertQA-MM | 77.3 | 64.4 | 80.7 | 68.7 | 71.0 |
  12. Grok 4.20 xAI65.8%as of 2026-04
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    publisher
    Meta
    locator
    Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Grok 4.2 Reasoning; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    reported by
    independent run
    published
    2026-04-08
    retrieved
    2026-09-07
    confidence
    verified
    MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
    also reported in Muse Spark Eval Methodology (model card, 65.8, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column Grok 4.2 Reasoning); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 65.8%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 6; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  13. Kimi K2.5 Moonshot AI65.3
    Qwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column)
    Qwen/Qwen3.5-397B-A17B model card · model card · first-party
    publisher
    Hugging Face
    locator
    Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
    reported by
    independent run
    published
    2026-02-16
    retrieved
    2026-09-07
    confidence
    verified
    MedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
  14. Claude Opus 4.6 Anthropic64.8%as of 2026-04
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    publisher
    Meta
    locator
    Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Opus 4.6 Max; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
    reported by
    independent run
    published
    2026-04-08
    retrieved
    2026-09-07
    confidence
    verified
    MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
    also reported in Muse Spark Eval Methodology (model card, 64.8, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column Opus 4.6 Max); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 64.8%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 7; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
  15. Gemma 4 12B Google48.7%as of 2026-04
    Google's Gemma 4 model card, Unified 12B; protocol not stated
    Gemma 4 model card · model card · first-party
    publisher
    Google
    locator
    Benchmark Results table, Vision section, row MedXPertQA MM, column Gemma 4 12B Unified; header row: | | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) |
    reported by
    vendor-reported
    published
    2026-04-02
    retrieved
    2026-09-07
    confidence
    verified
    | MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | \- |
    also reported in Gemma 4 Technical Report (paper, 48.7, Gemma 4 Technical Report p. 6, Table 6 (vision benchmarks, thinking), row MedXPertQA MM, column 12B); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 48.7%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 8; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)

Board source: Best AI for Healthcare & Medical: LLM Leaderboard

  1. Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    embedded JSON-LD FAQPage, answer to "Which AI is best for doctors and clinicians?"; the Score chart JSON-LD carries the same entry as {"label":"Claude Fable 5.1 (max with fallback)","score":56.35682697075397,"detailsUrl":"/models/claude-fable-5-1"}; chart header "23 of 72 models"; retrieved 2026-09-07
    reported by
    official leaderboard
    retrieved
    2026-09-08
    confidence
    verified
    Based on the Artificial Analysis Healthcare & Medical Index, the top-performing AI models for healthcare and medical work are currently Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (56), Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (52), and Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (52). Rankings are updated as new models are released.
  2. Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    embedded JSON-LD FAQPage, answer to "Which AI is best for doctors and clinicians?"; the Score chart JSON-LD carries the same entry as {"label":"Claude Fable 5 (with fallback)","score":52.01006901640588,"detailsUrl":"/models/claude-fable-5"}; chart header "23 of 72 models"; retrieved 2026-09-07
    reported by
    official leaderboard
    retrieved
    2026-09-08
    confidence
    verified
    Based on the Artificial Analysis Healthcare & Medical Index, the top-performing AI models for healthcare and medical work are currently Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (56), Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (52), and Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (52). Rankings are updated as new models are released.
  3. all underlying benchmarks run independently by Artificial Analysis
    Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    embedded JSON-LD FAQPage, answer to "Which AI is best for doctors and clinicians?"; the Score chart JSON-LD carries the same entry as {"label":"Claude Opus 5 (max)","score":51.66169193981395,"detailsUrl":"/models/claude-opus-5"}; chart header "23 of 72 models"; retrieved 2026-09-07
    reported by
    official leaderboard
    retrieved
    2026-09-08
    confidence
    verified
    Based on the Artificial Analysis Healthcare & Medical Index, the top-performing AI models for healthcare and medical work are currently Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (56), Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (52), and Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (52). Rankings are updated as new models are released.
  4. GPT-5.6 Sol (max) OpenAI44as of 2026-09
    Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 9 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
    reported by
    official leaderboard
    retrieved
    2026-09-08
    confidence
    verified
    {"label":"GPT-5.6 Sol (max)","score":43.7633355450504,"detailsUrl":"/models/gpt-5-6-sol"}
  5. Grok 4.6 (high) xAI44as of 2026-09
    Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 8 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
    reported by
    official leaderboard
    retrieved
    2026-09-08
    confidence
    verified
    {"label":"Grok 4.6 (high)","score":43.92324392614704,"detailsUrl":"/models/grok-4-6"}
  6. Kimi K3 (max) Moonshot AI44as of 2026-09
    Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 10 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
    reported by
    official leaderboard
    retrieved
    2026-09-08
    confidence
    verified
    {"label":"Kimi K3 (max)","score":43.62067857170797,"detailsUrl":"/models/kimi-k3"}
  7. GPT-5.6 Terra (max) OpenAI39as of 2026-09
    Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 12 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
    reported by
    official leaderboard
    retrieved
    2026-09-08
    confidence
    verified
    {"label":"GPT-5.6 Terra (max)","score":38.66673219960154,"detailsUrl":"/models/gpt-5-6-terra"}
  8. GPT-5.6 Luna (max) OpenAI34as of 2026-09
    Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 15 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
    reported by
    official leaderboard
    retrieved
    2026-09-08
    confidence
    verified
    {"label":"GPT-5.6 Luna (max)","score":34.39525059118443,"detailsUrl":"/models/gpt-5-6-luna"}
  9. Inkling (xhigh) Thinking Machines31as of 2026-09
    Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 18 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
    reported by
    official leaderboard
    retrieved
    2026-09-08
    confidence
    verified
    {"label":"Inkling","score":30.52028583441501,"detailsUrl":"/models/inkling"}
  10. MiniMax-M3 MiniMax31as of 2026-09
    Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 17 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
    reported by
    official leaderboard
    retrieved
    2026-09-08
    confidence
    verified
    {"label":"MiniMax-M3","score":30.552965002205703,"detailsUrl":"/models/minimax-m3"}
  11. Mistral Medium 3.5 Mistral17as of 2026-09
    Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    page data payload initialModels (RSC flight, JSON-escaped inside self.__next_f.push in the cached HTML), entry for slug "mistral-medium-3-5", rank 22 of 23; the Score chart JSON-LD lists only the top 20; chart header "23 of 72 models"; retrieved 2026-09-07
    reported by
    official leaderboard
    retrieved
    2026-09-08
    confidence
    verified
    \"slug\":\"mistral-medium-3-5\",\"name\":\"Mistral Medium 3.5\" ... \"headlineValue\":17.2489866350548
  12. gpt-oss-120b (high) OpenAI14as of 2026-09
    Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
    publisher
    Artificial Analysis
    locator
    page data payload initialModels (RSC flight, JSON-escaped inside self.__next_f.push in the cached HTML), entry for slug "gpt-oss-120b", rank 23 of 23; the Score chart JSON-LD lists only the top 20; chart header "23 of 72 models"; retrieved 2026-09-07
    reported by
    official leaderboard
    retrieved
    2026-09-08
    confidence
    verified
    \"slug\":\"gpt-oss-120b\",\"name\":\"gpt-oss-120b (high)\" ... \"headlineValue\":14.423032488856672

PhysicianBench

9 rows · pass@1 success rate %

Board source: PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · official page: arxiv.org

  1. GPT-5.5 OpenAI46.3 ± 1.2as of 2026-05
    pass@1; Pass^3 28.0
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    official leaderboard
    published
    2026-05-04
    retrieved
    2026-09-07
    confidence
    verified
    GPT-5.5 46.3 ± 1.2 57.4 28.0 41.9
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 46.3 ± 1.2, arXiv abs page (landing page for the PDF))
  2. Claude Opus 4.6 Anthropic31.7 ± 2.3as of 2026-05
    Pass^3 18.0
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    official leaderboard
    published
    2026-05-04
    retrieved
    2026-09-07
    confidence
    verified
    Claude Opus 4.6 31.7 ± 2.3 41.5 18.0 25.2
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 31.7 ± 2.3, arXiv abs page (landing page for the PDF))
  3. Claude Opus 4.7 Anthropic29.3 ± 2.5as of 2026-05
    Pass^3 18.0
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    official leaderboard
    published
    2026-05-04
    retrieved
    2026-09-07
    confidence
    verified
    Claude Opus 4.7 29.3 ± 2.5 37.9 18.0 16.2
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 29.3 ± 2.5, arXiv abs page (landing page for the PDF))
  4. GPT-5.4 OpenAI27.7 ± 1.5as of 2026-05
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    official leaderboard
    published
    2026-05-04
    retrieved
    2026-09-07
    confidence
    verified
    GPT-5.4 27.7 ± 1.5 37.7 13.0 39.8
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 27.7 ± 1.5, arXiv abs page (landing page for the PDF))
  5. Claude Sonnet 4.6 Anthropic23.0 ± 2.6as of 2026-05
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    official leaderboard
    published
    2026-05-04
    retrieved
    2026-09-07
    confidence
    verified
    Claude Sonnet 4.6 23.0 ± 2.6 33.2 9.0 22.3
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 23.0 ± 2.6, arXiv abs page (landing page for the PDF))
  6. Kimi-K2.6 Moonshot AI17.0 ± 2.6as of 2026-05
    open source
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    official leaderboard
    published
    2026-05-04
    retrieved
    2026-09-07
    confidence
    verified
    Kimi-K2.6 17.0 ± 2.6 26.3 5.0 42.4
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 17.0 ± 2.6, arXiv abs page (landing page for the PDF))
  7. Qwen3.6-Plus Alibaba13.7 ± 4.0as of 2026-05
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    official leaderboard
    published
    2026-05-04
    retrieved
    2026-09-07
    confidence
    verified
    Qwen3.6-Plus 13.7 ± 4.0 22.6 2.0 28.0
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 13.7 ± 4.0, arXiv abs page (landing page for the PDF))
  8. Gemini Pro 3.1 Google6.0 ± 1.0as of 2026-05
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
    reported by
    official leaderboard
    published
    2026-05-04
    retrieved
    2026-09-07
    confidence
    verified
    Gemini Pro 3.1 6.0 ± 1.0 9.3 3.0 30.4
    also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 6.0 ± 1.0, arXiv abs page (landing page for the PDF))
  9. Grok-4.20 xAI5.3 ± 3.2as of 2026-05
    publisher
    Stanford University (HealthRex; Liu, Chen et al.)
    locator
    p. 8, Table 2 (Proprietary Models block)
    reported by
    official leaderboard
    published
    2026-05-04
    retrieved
    2026-09-07
    confidence
    verified
    Grok-4.20 5.3 ± 3.2 9.7 1.0 16.7

EHR-Complex

6 rows · exact-match accuracy

Board source: EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · official page: arxiv.org

  1. GPT-5.4 (high reasoning) OpenAI0.65as of 2026-06
    average over 12 intent columns; run as human-validation configuration, not in the headline 12-model table
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 15, Table 10 (Strong commercial model results), Avg. column
    reported by
    official leaderboard
    published
    2026-06-22
    retrieved
    2026-09-07
    confidence
    verified
    GPT-5.4 (high) 0.85 0.36 0.78 0.34 0.8 0.37 0.93 0.76 0.91 0.49 0.84 0.36 0.65
    also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.65, arXiv abs page (landing page for the PDF))
  2. Gemini 3.1 Pro Google0.63as of 2026-06
    validation configuration
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 15, Table 10 (Strong commercial model results), Avg. column
    reported by
    official leaderboard
    published
    2026-06-22
    retrieved
    2026-09-07
    confidence
    verified
    Gemini 3.1 Pro 0.87 0.32 0.67 0.38 0.84 0.37 0.92 0.63 0.92 0.44 0.83 0.34 0.63
    also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.63, arXiv abs page (landing page for the PDF))
  3. Kimi-K2.5 Moonshot AI0.62as of 2026-06
    headline 12-model evaluation, top open-weight
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. column
    reported by
    official leaderboard
    published
    2026-06-22
    retrieved
    2026-09-07
    confidence
    verified
    Kimi-K2.5 0.89 0.34 0.8 0.27 0.73 0.35 0.92 0.72 0.89 0.42 0.84 0.29 0.62
    also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.62, arXiv abs page (landing page for the PDF))
  4. Qwen3.5-397B Alibaba0.62as of 2026-06
    headline evaluation
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. column
    reported by
    official leaderboard
    published
    2026-06-22
    retrieved
    2026-09-07
    confidence
    verified
    Qwen3.5-397B 0.88 0.39 0.72 0.3 0.8 0.38 0.91 0.63 0.91 0.45 0.78 0.32 0.62
    also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.62, arXiv abs page (landing page for the PDF))
  5. GPT-5.4 (low reasoning) OpenAI0.58as of 2026-06
    validation configuration
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 15, Table 10 (Strong commercial model results), Avg. column
    reported by
    official leaderboard
    published
    2026-06-22
    retrieved
    2026-09-07
    confidence
    verified
    GPT-5.4 (low) 0.81 0.32 0.65 0.26 0.73 0.31 0.88 0.66 0.86 0.43 0.8 0.29 0.58
    also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.58, arXiv abs page (landing page for the PDF))
  6. Claude Sonnet 4.6 Anthropic0.36as of 2026-06
    validation configuration
    publisher
    Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
    locator
    p. 15, Table 10 (Strong commercial model results), Avg. column
    reported by
    official leaderboard
    published
    2026-06-22
    retrieved
    2026-09-07
    confidence
    verified
    Claude Sonnet 4.6 0.59 0.21 0.45 0.12 0.45 0.12 0.83 0.2 0.49 0.33 0.4 0.11 0.36
    also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.36, arXiv abs page (landing page for the PDF))

WHBench

6 rows · mean normalized percentage 0-100

Board source: WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · official page: arxiv.org

  1. Claude Opus 4.6 Anthropic72.1%as of 2026-03
    95% CI 69.6-74.4; evaluations run March 2026
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    official leaderboard
    published
    2026-07-23
    retrieved
    2026-09-07
    confidence
    verified
    1 Claude Opus 4.6 72.1 [69.6, 74.4] 35.5 58.2 6.4 12.8
    also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 72.1%, same paper, abstract page)
  2. Claude Sonnet 4.6 Anthropic67.1%as of 2026-03
    95% CI 64.5-69.6
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    official leaderboard
    published
    2026-07-23
    retrieved
    2026-09-07
    confidence
    verified
    2 Claude Sonnet 4.6 67.1 [64.5, 69.6] 22.7 67.4 9.9 27.0
    also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 67.1%, same paper, abstract page)
  3. GPT-5.4 OpenAI66.8%as of 2026-03
    95% CI 64.5-69.2
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    official leaderboard
    published
    2026-07-23
    retrieved
    2026-09-07
    confidence
    verified
    3 GPT-5.4 66.8 [64.5, 69.2] 21.3 67.4 11.3 47.5
    also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 66.8%, same paper, abstract page)
  4. Gemini 3 Flash Preview Google64.7%as of 2026-03
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    official leaderboard
    published
    2026-07-23
    retrieved
    2026-09-07
    confidence
    verified
    4 Gemini 3 Flash Preview 64.7 [61.7, 67.7] 25.5 62.4 12.1 32.6
    also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 64.7%, same paper, abstract page)
  5. GPT-4.1 OpenAI51.8%as of 2026-03
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    official leaderboard
    published
    2026-07-23
    retrieved
    2026-09-07
    confidence
    verified
    11 GPT-4.1 51.8 [49.2, 54.3] 3.5 61.7 34.8 61.0
    also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 51.8%, same paper, abstract page)
  6. GPT-4o OpenAI44.6%as of 2026-03
    publisher
    Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
    locator
    p. 5, Table 3 (WHBench v3.0 leaderboard)
    reported by
    official leaderboard
    published
    2026-07-23
    retrieved
    2026-09-07
    confidence
    verified
    16 GPT-4o 44.6 [41.8, 47.4] 1.4 42.5 56.0 83.7
    also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 44.6%, same paper, abstract page)

HealthAdminBench

7 rows · percentage end-to-end task success 0-100

Board source: HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) · official page: healthadminbench.stanford.edu · paper: arxiv.org

  1. Claude Opus 4.6 (computer-use agent) Anthropic36.3%as of 2026-04
    screenshot-only, detailed prompting; native CUA harness; subtask rate 78.4%
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    official leaderboard
    published
    2026-04-10
    retrieved
    2026-09-07
    confidence
    verified
    Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate
    also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 36.3%, Section 'LLMs struggle with long-horizon tasks')
  2. GPT-5.4 (computer-use agent) OpenAI26.7%as of 2026-04
    screenshot-only, detailed prompting; subtask rate 82.8%
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    official leaderboard
    published
    2026-04-10
    retrieved
    2026-09-07
    confidence
    verified
    Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate
    also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 26.7%, Section 'LLMs struggle with long-horizon tasks')
  3. Kimi K2.5 Moonshot AI15.6%as of 2026-04
    screenshot-only, detailed prompting
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    official leaderboard
    published
    2026-04-10
    retrieved
    2026-09-07
    confidence
    verified
    Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate
    also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 15.6%, Section 'LLMs struggle with long-horizon tasks')
  4. Claude Opus 4.6 (standardized harness) Anthropic14.8%as of 2026-04
    screenshot-only, detailed prompting; authors' standardized harness, no native CUA
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    official leaderboard
    published
    2026-04-10
    retrieved
    2026-09-07
    confidence
    verified
    Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate
  5. Qwen 3.5 Alibaba13.3%as of 2026-04
    screenshot-only, detailed prompting
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    official leaderboard
    published
    2026-04-10
    retrieved
    2026-09-07
    confidence
    verified
    Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate
    also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 13.3%, Section 'LLMs struggle with long-horizon tasks')
  6. Gemini 3.1 Pro Google11.9%as of 2026-04
    screenshot-only, detailed prompting
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    official leaderboard
    published
    2026-04-10
    retrieved
    2026-09-07
    confidence
    verified
    Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate
    also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 11.9%, Section 'LLMs struggle with long-horizon tasks')
  7. GPT-5.4 (standardized harness) OpenAI5.9%as of 2026-04
    screenshot-only, detailed prompting; authors' standardized harness, no native CUA
    publisher
    Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
    locator
    p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
    reported by
    official leaderboard
    published
    2026-04-10
    retrieved
    2026-09-07
    confidence
    verified
    Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate

OpenAI Dynamic Mental Health Evaluations

15 rows · compliance rate per metric

Board source: GPT-5.6 - August Updates (system card addendum)

  1. GPT-6 Astra OpenAI1.000as of 2026-09
    dynamic multi-turn adversarial user simulations, metric 'safe' (share of assistant messages that do not violate safety policies): mental health 1.000, emotional reliance 0.993, self-harm 0.989.
    GPT-6 Astra System Card · system card · first-party
    publisher
    OpenAI
    locator
    p. 20, Table 7 (Dynamic Benchmarks with Adversarial User Simulations), column 'gpt-6 Astra'
    reported by
    vendor-reported
    published
    2026-09-03
    retrieved
    2026-09-08
    confidence
    verified
    Mental health 0.914 0.820 0.991 1.000
  2. GPT-5.5 Instant OpenAI0.999as of 2026-05
    mental health 0.999, emotional reliance 0.963, self-harm 0.913
    GPT-5.5 Instant System Card (PDF) · system card · first-party
    publisher
    OpenAI
    locator
    p. 6, Table 3, sec. 3.3
    reported by
    vendor-reported
    published
    2026-05-05
    retrieved
    2026-09-07
    confidence
    verified
    Mental health 0.818 1.000 1.000 0.999 / Emotional reliance 0.976 0.992 0.995 0.963 / Self-harm 0.842 0.976 0.924 0.913 (gpt-5.5-instant column)
    also reported in GPT-5.6 - August Updates (system card addendum) (system card, 0.996, conflicting, p. 9, sec. 3.4)
  3. GPT-5.5 Instant (June Update) OpenAI0.991as of 2026-08
    mental health 0.991, emotional reliance 0.989, self-harm 0.967; measured at lowest reasoning deployment settings
    publisher
    OpenAI
    locator
    p. 9, table "Dynamic Benchmarks with Adversarial User Simulations", sec. 3.4
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-07
    confidence
    verified
    Mental health 0.984 0.996 0.981 0.991 0.981 0.977 / Emotional reliance 0.987 0.965 0.988 0.989 0.961 0.965 / Self-harm 0.816 0.839 0.852 0.967 0.901 0.911 (GPT-5.5 Instant (June Update) column)
  4. GPT-5.6 Sol (July) OpenAI0.991as of 2026-07
    July system card version (API, Codex and Work deployments), dynamic multi-turn adversarial user simulations, metric not_unsafe: mental health 0.991, emotional reliance 0.953, self-harm 0.856. Distinct from the August ChatGPT production version measured in row 193 (0.981 / 0.961 / 0.901).
    GPT-5.6 System Card (PDF) · system card · first-party
    publisher
    OpenAI
    locator
    p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    Mental health 0.753 0.975 0.985 0.820 0.991 0.985 0.989 / Emotional reliance 0.857 0.953 0.985 0.915 0.953 0.976 0.957 / Self-harm 0.904 0.955 0.977 0.868 0.856 0.947 0.905 (gpt-5.6-sol column)
  5. GPT-5.6 Luna (July) OpenAI0.989as of 2026-07
    July system card version, dynamic multi-turn adversarial user simulations, metric not_unsafe: mental health 0.989, emotional reliance 0.957, self-harm 0.905. Distinct from the August ChatGPT production version measured in row 194 (0.977 / 0.965 / 0.911).
    GPT-5.6 System Card (PDF) · system card · first-party
    publisher
    OpenAI
    locator
    p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    Mental health ... 0.989 / Emotional reliance ... 0.957 / Self-harm ... 0.905 (gpt-5.6-luna column)
  6. GPT-5.3 Instant OpenAI0.985as of 2026-03
    mental health 0.985, emotional reliance 0.992, self-harm 0.911; dynamic multi-turn adversarial user simulations, metric not_unsafe
    GPT-5.3 Instant System Card · system card · first-party
    publisher
    OpenAI
    locator
    p. 2, Table 2
    reported by
    vendor-reported
    published
    2026-03-03
    retrieved
    2026-09-07
    confidence
    verified
    mental health* 0.832 1.000 0.985 / emotional reliance* 0.945 0.952 0.992 / Self-harm * 0.845 0.920 0.911
    also reported in GPT-5.5 Instant System Card (PDF) (system card, 1.000, conflicting, p. 6, Table 3, sec. 3.3); GPT-5.6 - August Updates (system card addendum) (system card, 0.984, conflicting, p. 9, sec. 3.4)
  7. GPT-5.4 mini OpenAI0.985as of 2026-03
    mental health 0.985, emotional reliance 0.977, self-harm 0.980; dynamic multi-turn adversarial user simulations, metric not_unsafe
    GPT-5.4 Thinking System Card · system card · first-party
    publisher
    OpenAI
    locator
    p. 32, Table 17 (Appendix 6.1, added 2026-03-17)
    reported by
    vendor-reported
    published
    2026-03-05
    retrieved
    2026-09-07
    confidence
    verified
    Mental health 0.753 0.975 0.985 0.985 / Emotional reliance 0.857 0.953 0.985 0.977 / Self-harm 0.904 0.955 0.977 0.980
  8. GPT-5.4 Thinking OpenAI0.985as of 2026-03
    mental health 0.985, emotional reliance 0.985, self-harm 0.977; dynamic multi-turn adversarial user simulations, metric not_unsafe
    GPT-5.4 Thinking System Card · system card · first-party
    publisher
    OpenAI
    locator
    p. 32, Table 17 (same values as p. 5, Table 2)
    reported by
    vendor-reported
    published
    2026-03-05
    retrieved
    2026-09-07
    confidence
    verified
    Mental health 0.753 0.975 0.985 0.985 / Emotional reliance 0.857 0.953 0.985 0.977 / Self-harm 0.904 0.955 0.977 0.980
    also reported in GPT-5.5 System Card (system card, 0.985, p. 12, Table 8 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2); GPT-5.6 System Card (PDF) (system card, 0.985, p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2)
  9. GPT-5.6 Terra OpenAI0.985as of 2026-07
    mental health 0.985, emotional reliance 0.976, self-harm 0.947
    GPT-5.6 System Card (PDF) · system card · first-party
    publisher
    OpenAI
    locator
    p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2
    reported by
    vendor-reported
    published
    2026-07-09
    retrieved
    2026-09-07
    confidence
    verified
    Mental health ... 0.985 / Emotional reliance ... 0.976 / Self-harm ... 0.947 (gpt-5.6-terra column)
  10. GPT-5.5 OpenAI0.981as of 2026-04
    mental health 0.981, emotional reliance 0.981, self-harm 0.937
    GPT-5.5 System Card · system card · first-party
    publisher
    OpenAI
    locator
    p. 12, Table 8 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2
    reported by
    vendor-reported
    published
    2026-04-23
    retrieved
    2026-09-07
    confidence
    verified
    Mental health 0.753 0.975 0.985 0.981 / Emotional reliance 0.857 0.953 0.985 0.981 / Self-harm 0.904 0.955 0.977 0.937 (gpt-5.5 column)
    also reported in GPT-5.6 System Card (PDF) (system card, 0.820, conflicting, p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2)
  11. GPT-5.5 Instant (May Update) OpenAI0.981as of 2026-08
    mental health 0.981, emotional reliance 0.988, self-harm 0.852
    publisher
    OpenAI
    locator
    p. 9, table "Dynamic Benchmarks with Adversarial User Simulations", sec. 3.4
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-07
    confidence
    verified
    Mental health 0.984 0.996 0.981 0.991 0.981 0.977 / Emotional reliance 0.987 0.965 0.988 0.989 0.961 0.965 / Self-harm 0.816 0.839 0.852 0.967 0.901 0.911 (GPT-5.5 Instant (May Update) column)
  12. GPT-5.6 Sol (August) OpenAI0.981as of 2026-08
    mental health 0.981, emotional reliance 0.961, self-harm 0.901; OpenAI flags statistically significant offline self-harm regression vs GPT-5.5 June, not reproduced online
    publisher
    OpenAI
    locator
    p. 9, table "Dynamic Benchmarks with Adversarial User Simulations", sec. 3.4
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-07
    confidence
    verified
    Mental health 0.984 0.996 0.981 0.991 0.981 0.977 / Emotional reliance 0.987 0.965 0.988 0.989 0.961 0.965 / Self-harm 0.816 0.839 0.852 0.967 0.901 0.911 (GPT-5.6 Sol (August) column)
  13. GPT-5.6 Luna (August) OpenAI0.977as of 2026-08
    mental health 0.977, emotional reliance 0.965, self-harm 0.911
    publisher
    OpenAI
    locator
    p. 9, table "Dynamic Benchmarks with Adversarial User Simulations", sec. 3.4
    reported by
    vendor-reported
    published
    2026-08-06
    retrieved
    2026-09-07
    confidence
    verified
    Mental health 0.984 0.996 0.981 0.991 0.981 0.977 / Emotional reliance 0.987 0.965 0.988 0.989 0.961 0.965 / Self-harm 0.816 0.839 0.852 0.967 0.901 0.911 (GPT-5.6 Luna (August) column)
  14. GPT-5.2 Thinking OpenAI0.975as of 2026-03
    mental health 0.975, emotional reliance 0.953, self-harm 0.955; dynamic multi-turn adversarial user simulations, metric not_unsafe
    GPT-5.4 Thinking System Card · system card · first-party
    publisher
    OpenAI
    locator
    p. 32, Table 17
    reported by
    vendor-reported
    published
    2026-03-05
    retrieved
    2026-09-07
    confidence
    verified
    Mental health 0.753 0.975 0.985 0.985 / Emotional reliance 0.857 0.953 0.985 0.977 / Self-harm 0.904 0.955 0.977 0.980
    also reported in GPT-5.5 System Card (system card, 0.975, p. 12, Table 8 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2); GPT-5.6 System Card (PDF) (system card, 0.975, p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2)
  15. GPT-5.1 Thinking OpenAI0.753as of 2026-03
    mental health 0.753, emotional reliance 0.857, self-harm 0.904; dynamic multi-turn adversarial user simulations, metric not_unsafe
    GPT-5.4 Thinking System Card · system card · first-party
    publisher
    OpenAI
    locator
    p. 32, Table 17 (columns gpt-5.1-thinking, gpt-5.2-thinking, gpt-5.4-thinking, gpt-5.4-mini)
    reported by
    vendor-reported
    published
    2026-03-05
    retrieved
    2026-09-07
    confidence
    verified
    Mental health 0.753 0.975 0.985 0.985 / Emotional reliance 0.857 0.953 0.985 0.977 / Self-harm 0.904 0.955 0.977 0.980
    also reported in GPT-5.5 System Card (system card, 0.753, p. 12, Table 8 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2); GPT-5.6 System Card (PDF) (system card, 0.753, p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2)

Documents on file

The distinct documents the sections above draw on, with the number of index rows each one backs.

documentkindrows
GPT-5.6 Preview System Card
deploymentsafety.openai.com
system card · first-party22
GPT-5.6 System Card
deploymentsafety.openai.com
system card · first-party21
GPT-5.6 - August Updates (system card addendum)
cdn.openai.com
system card · first-party16
Health Optimization Bench: Arcophos harness run
healthoptimizationbench.com
Arcophos run · first-party16
Best AI for Healthcare & Medical: LLM Leaderboard
artificialanalysis.ai
official leaderboard · first-party12
CHI-Bench leaderboard (actAVA)
actava.ai
official leaderboard · first-party12
MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results)
arise-ai.org
official leaderboard · first-party12
Vals AI MedScribe leaderboard
vals.ai
official leaderboard · first-party12
HealthAgentBench detailed results
microsoft.github.io
official leaderboard · first-party11
HealthAgentBench leaderboard
microsoft.github.io
official leaderboard · first-party11
Vals AI MedCode leaderboard
vals.ai
official leaderboard · first-party11
MedHELM leaderboard (medhelm.org), v5.0.0
medhelm.org
official leaderboard · first-party10
MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate)
leaderboard.medhelm.org
official leaderboard · first-party10
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1)
arxiv.org
paper9
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF)
arxiv.org
paper9
MAST: Medical AI Superintelligence Test leaderboard (General board)
arise-ai.org
official leaderboard · first-party8
MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai
benchlm.ai
mirror8
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240)
arxiv.org
paper8
System Card: Claude Opus 5
anthropic.com
system card · first-party8
GPT-5.6 System Card (PDF)
deploymentsafety.openai.com
system card · first-party7
HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1)
arxiv.org
paper7
Claude Fable 5.1 and Claude Mythos 5.1 System Card
anthropic.com
system card · first-party6
EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301)
arxiv.org
paper6
EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF)
arxiv.org
paper6
Introducing Muse Spark: Scaling Towards Personal Superintelligence
ai.meta.com
launch post · first-party6
Muse Spark Eval Methodology
ai.meta.com
model card · first-party6
WHBench (arXiv 2604.00024 abstract page)
arxiv.org
paper6
WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2)
arxiv.org
paper6
GPT-5.5 Instant System Card
deploymentsafety.openai.com
system card · first-party5
Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog)
kineticsystems.ai
vendor post5
Qwen3.8-Max: A New Bar for Coding and Cowork
qwen.ai
launch post · first-party5
Baichuan-M3 Technical Report
arxiv.org
paper4
Claude Fable 5 and Claude Mythos 5 System Card
anthropic.com
system card · first-party4
GPT-5.4 Thinking System Card
deploymentsafety.openai.com
system card · first-party4
GPT-5.5 System Card
deploymentsafety.openai.com
system card · first-party4
GPT-6 Astra System Card
deploymentsafety.openai.com
system card · first-party4
gpt-oss-120b & gpt-oss-20b Model Card
deploymentsafety.openai.com
model card · first-party4
GPT-6 Astra System Card - HealthBench (Deployment Safety Hub)
deploymentsafety.openai.com
system card · first-party3
Qwen/Qwen3.5-397B-A17B model card
huggingface.co
model card · first-party3
System Card: Claude Opus 4.8
anthropic.com
system card · first-party3
System Card: Claude Sonnet 5
anthropic.com
system card · first-party3
GPT-5.3 Instant System Card
deploymentsafety.openai.com
system card · first-party2
GPT-5.5 Instant System Card (PDF)
deploymentsafety.openai.com
system card · first-party2
GPT-5.5 System Card
deploymentsafety.openai.com
system card · first-party2
Muse Spark 1.1 Evaluation Report
research.meta.ai
model card · first-party2
Qwen3.7-Plus: Multimodal Agent Intelligence
qwen.ai
launch post · first-party2
Gemma 4 model card
ai.google.dev
model card · first-party1
Gemma 4 Technical Report
arxiv.org
paper1
GPT-5 System Card
cdn.openai.com
system card · first-party1
GPT-5.3 Instant System Card
deploymentsafety.openai.com
system card · first-party1
GPT-6 Astra: A new generation of intelligence
openai.com
launch post · first-party1
MAI-Thinking-1: Building a Hill-Climbing Machine
microsoft.ai
model card · first-party1

Corrections go through the same route as everything else here: a better document replaces a weaker one, the row's confidence moves, and the change is dated on the updates page. The full record, sources included, is downloadable from the data page.