Sources
218 of 218 rows documented · 52 documents, 40 first-party · updated September 8, 2026
Every number on this index was copied out of a document, and this page is the list of those documents. A row records where its score was read: the publication itself, the passage or table it sits in when that has been captured, who published it, when the index retrieved it, and how far the reading has been checked. First-party documents rank highest, meaning a vendor's system card for a vendor-reported number or the maintainers' own board for a leaderboard number; aggregator pages that repeat a figure are kept as corroboration and never stand in for the original. Rows whose document has not been located yet say so instead of disappearing.
Better citations do not make the boards comparable. Each benchmark below keeps its own scale, grader and task set, so a score on one section says nothing about a score on the next, and no number on this page should be lined up against a number from another section. The comparability rules are on the methodology page.
Reading the confidence field
- verified: the number was read from the document at the locator given
- partial: the document is identified, the exact passage is not yet on file
- unverified: no document has been checked for this number yet
- disputed: documents on file disagree about this number
HealthBench Professional
22 rows · 0 to 1Board source: healthbenchprofessional.com · paper: arxiv.org · full board: healthbenchprofessional.com
- Claude Fable 5 Anthropic0.660as of 2026-06length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 70.3%). Measured as Claude Mythos 5; Anthropic itself prints 66.0 in the Fable 5 column of the Opus 5 card with footnote Mythos 5.Claude Fable 5 and Claude Mythos 5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- p. 252, Table 8.1.A, row HealthBench Professional, column Mythos 5 (Fable 5 column '-'); Figure 8.18.2.A p. 298 bar label 66.0% on Claude Mythos 5
- reported by
- vendor-reported
- published
- 2026-06-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional 66.0 - 64.7 56.9 51.8 -
also reported in System Card: Claude Opus 5 (system card, 66.0, p. 152, Table 8.1.A, column Fable 5 (footnote 11: Mythos 5)); Claude Fable 5.1 and Claude Mythos 5.1 System Card (system card, 63.3%, conflicting, p. 167, Table 8.1.A, column 'Claude Fable 5/Mythos 5') - GPT-6 Astra OpenAI0.634as of 2026-09length-adjusted, max reasoning effort (69.5 unadjusted, 4,097 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'.GPT-6 Astra System Card · system card · first-party
- publisher
- OpenAI
- locator
- p. 19, sec. 6.1 (Table 6 cell, column 'gpt-6 Astra': 63.4 (69.5, 4097))
- reported by
- vendor-reported
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
Astra has a length-adjusted HealthBench Professional score of 63.4 (+2.9 relative to GPT-5.6 Sol), HealthBench score of 58.1 (+1.1), HealthBench Hard score of 36.3 (+3.2), and HealthBench Consensus score of 95.8 (+0.3).
also reported in GPT-6 Astra: A new generation of intelligence (launch post, 63.4%, Science and Health table); GPT-6 Astra System Card - HealthBench (Deployment Safety Hub) (system card, 63.4, mirror, section 6.1, HTML rendering of the same card) - Claude Fable 5.1 Anthropic0.621as of 2026-09length-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%).Claude Fable 5.1 and Claude Mythos 5.1 System Card · system card · first-party
- publisher
- Anthropic
- locator
- p. 199, sec. 8.17.2 (same figure printed as 62.1% in Table 8.1.A, p. 167)
- reported by
- vendor-reported
- published
- 2026-09-01
- retrieved
- 2026-09-07
- confidence
- verified
On HealthBench Professional, Claude Fable 5.1 achieved a raw score of 74.2%, ahead of Claude Opus 5 at 73.4%, Fable 5 at 68.9%, and Claude Sonnet 5 at 62.4%. After length adjustment, which penalizes verbose model responses, Fable 5.1 achieved a score of 62.1%.
- GPT-5.6 Sol OpenAI0.605as of 2026-06length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
also reported in GPT-5.6 Preview System Card (system card, 60.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL) - Claude Opus 5 Anthropic0.598as of 2026-07length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%).System Card: Claude Opus 5 · system card · first-party
- publisher
- Anthropic
- locator
- p. 189, section 8.15.2 HealthBench Professional results; also Table 8.1.A p. 152 ('HealthBench Professional 59.8 ...')
- reported by
- vendor-reported
- published
- 2026-07-24
- retrieved
- 2026-09-07
- confidence
- verified
Claude Opus 5 achieved a raw score of 73.4%, which is the highest amongst all Claude models, ahead of Claude Mythos 5 at 70.3%, Claude Opus 4.8 at 60.3%, and Claude Sonnet 5 at 62.4%. After length adjustment, which penalizes verbose model responses, Claude Opus 5 achieved a score of 59.8%.
also reported in Claude Fable 5.1 and Claude Mythos 5.1 System Card (system card, 59.8%, p. 167, Table 8.1.A, column 'Claude Opus 5'; raw 73.4% also reprinted on p. 199, sec. 8.17.2) - Muse Spark 1.1 Meta0.593as of 2026-07length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44)Muse Spark 1.1 Evaluation Report · model card · first-party
- publisher
- Meta
- locator
- p. 101, Figure 44 'General capability benchmark results' (image), row HealthBench Professional, column Muse Spark 1.1; protocol p. 104 (printed 103): HealthBench Pro comprises 525 evaluation data points graded by rubrics. We use GPT-5.4 with low reasoning effort as the grader and report the length-normalized rubric score as done in their paper.
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
- Claude Sonnet 5 Anthropic0.578as of 2026-06length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).System Card: Claude Sonnet 5 · system card · first-party
- publisher
- Anthropic
- locator
- p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 5; Figure 8.12.2.A p. 139
- reported by
- vendor-reported
- published
- 2026-06-30
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional 57.8 44.2 51.8 -
also reported in System Card: Claude Opus 5 (system card, 57.8%, p. 189, section 8.15.2, Figure 8.15.2.A) - GPT-5.6 Terra OpenAI0.577as of 2026-06length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
also reported in GPT-5.6 Preview System Card (system card, 57.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA) - Claude Opus 4.8 Anthropic0.558as of 2026-05length-adjusted, adaptive thinking at max effort, Claude Sonnet 4.6 grader (the grader used by the Opus 4.8 card itself); the other Claude rows on this board use the Claude Opus 4.8 grader, under which Anthropic later prints 57.4 for this model.System Card: Claude Opus 4.8 · system card · first-party
- publisher
- Anthropic
- locator
- p. 228, section 8.14.1 HealthBench Professional; Figure 8.14.A p. 229
- reported by
- vendor-reported
- published
- 2026-05-28
- retrieved
- 2026-09-07
- confidence
- verified
Claude Opus 4.8 scores 55.8%, a meaningful improvement over Claude Opus 4.7 at 51.9% and Claude Sonnet 4.6 at 41.7%.
also reported in Claude Fable 5 and Claude Mythos 5 System Card (system card, 56.9, conflicting, p. 252, Table 8.1.A, column Opus 4.8); System Card: Claude Opus 5 (system card, 57.4, conflicting, p. 152, Table 8.1.A, column Opus 4.8) - GPT-5.6 Luna OpenAI0.557as of 2026-06length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
also reported in GPT-5.6 Preview System Card (system card, 55.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA) - Muse Spark Meta0.541as of 2026-07length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44Muse Spark 1.1 Evaluation Report · model card · first-party
- publisher
- Meta
- locator
- p. 101, Figure 44 (image), row HealthBench Professional, column Muse Spark; protocol p. 104 (printed 103)
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
- GPT-5.6 Sol (August) OpenAI0.540as of 2026-08ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars)GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional 32.9 (33.8, 2,285) 38.4 (40.7, 2,775) 54.0 (56.6, 2,894) 44.1 (46.8, 2,920)
- Claude Opus 4.7 Anthropic0.519as of 2026-05length-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 trials; comparison model in the Opus 4.8 cardSystem Card: Claude Opus 4.8 · system card · first-party
- publisher
- Anthropic
- locator
- p. 228, section 8.14.1 HealthBench Professional; Figure 8.14.A p. 229
- reported by
- vendor-reported
- published
- 2026-05-28
- retrieved
- 2026-09-07
- confidence
- verified
Claude Opus 4.8 scores 55.8%, a meaningful improvement over Claude Opus 4.7 at 51.9% and Claude Sonnet 4.6 at 41.7%.
- GPT-5.5 OpenAI0.518as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (57.2 unadjusted, 3818 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
also reported in GPT-5.6 Preview System Card (system card, 51.8, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.5) - GPT-5.4 OpenAI0.481as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.4
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
also reported in GPT-5.6 Preview System Card (system card, 48.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.4) - GPT-5 OpenAI0.462as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
also reported in GPT-5.6 Preview System Card (system card, 46.2, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5) - GPT-5.2 OpenAI0.459as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
also reported in GPT-5.6 Preview System Card (system card, 45.9, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.2) - Claude Sonnet 4.6 Anthropic0.442as of 2026-06length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Sonnet 5 card. Other Anthropic prints: 44.4% (Fable card Figure 8.18.2.A), 41.7% (Opus 4.8 card, Sonnet 4.6 grader)System Card: Claude Sonnet 5 · system card · first-party
- publisher
- Anthropic
- locator
- p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 4.6
- reported by
- vendor-reported
- published
- 2026-06-30
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional 57.8 44.2 51.8 -
also reported in System Card: Claude Opus 4.8 (system card, 41.7%, conflicting, p. 228, section 8.14.1) - GPT-5.6 Luna (August) OpenAI0.441as of 2026-08ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars)GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional 32.9 (33.8, 2,285) 38.4 (40.7, 2,775) 54.0 (56.6, 2,894) 44.1 (46.8, 2,920)
- GPT-5.1 OpenAI0.396as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.1
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
also reported in GPT-5.6 Preview System Card (system card, 39.6, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.1) - GPT-5.5 Instant OpenAI0.384as of 2026-05length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.GPT-5.5 Instant System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT
- reported by
- vendor-reported
- published
- 2026-05-05
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Professional 37.6 (40.4, 2,973) 35.7 (38.3, 2,872) 32.9 (33.8, 2,285) 38.4 (40.7, 2,775)
also reported in GPT-5.6 - August Updates (system card addendum) (system card, 38.4, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant) - MAI-Thinking-1 Microsoft0.350as of 2026-08length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision.MAI-Thinking-1: Building a Hill-Climbing Machine · model card · first-party
- publisher
- Microsoft AI
- locator
- p. 54, Table 12 'Post-trained model evaluation results on various public benchmarks', Health group, column HealthBench Prof.; protocol Appendix K.6 p. 106: HealthBench Professional introduces a length penalty for the primary metric, to correct for a well-observed correlation between lengthy responses and artificially increased LLM-grader scores. For all reported scores, we use the standard GPT-5.4 grader and rubrics provided by OpenAI.
- reported by
- vendor-reported
- published
- 2026-08-12
- retrieved
- 2026-09-07
- confidence
- verified
Model AIR-Bench CyberSec Instruct CyberSec Auto Long Fact Truthful QA HealthBench Prof. MedXpert QA MAI-Thinking-1 88 63 63 98 88 35 43 Sonnet 4.6 88 62 56 98 88 38 49
HealthBench Hard
16 rows · 0 to 1Board source: healthbenchhard.ai · paper: arxiv.org · full board: healthbenchhard.ai
- Muse Spark Meta0.428as of 2026-04raw score (no length adjustment), GPT-4.1 grader via the OpenAI simple-evals implementation, Muse Spark Thinking; Meta run, launch-post benchmark table.Introducing Muse Spark: Scaling Towards Personal Superintelligence · launch post · first-party
- publisher
- Meta
- locator
- Launch-post benchmark table image, HEALTH section, row HealthBench Hard, column Muse Spark Thinking; identical table in the Eval Methodology PDF p. 5; protocol p. 2: HealthBench Hard: This is a subset of OpenAI's HealthBench benchmark, containing 1000 prompts. We used the same implementation as in the OpenAI’s official simple-evals repo, with GPT-4.1-genai as the LLM-as-judge model.
- reported by
- vendor-reported
- published
- 2026-04-08
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard | Open-Ended Health Queries | 42.8 | 14.8 | 20.6 | 40.1 | 20.3 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT-5.4 Xhigh, Grok 4.2 Reasoning)
also reported in Muse Spark Eval Methodology (model card, 42.8, p. 5 results table image, row HealthBench Hard; protocol p. 2) - GPT-6 Astra OpenAI0.363as of 2026-09length-adjusted, max reasoning effort (37.8 unadjusted, 2,192 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'.GPT-6 Astra System Card · system card · first-party
- publisher
- OpenAI
- locator
- p. 19, sec. 6.1 (Table 6 cell, column 'gpt-6 Astra': 36.3 (37.8, 2192))
- reported by
- vendor-reported
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
Astra has a length-adjusted HealthBench Professional score of 63.4 (+2.9 relative to GPT-5.6 Sol), HealthBench score of 58.1 (+1.1), HealthBench Hard score of 36.3 (+3.2), and HealthBench Consensus score of 95.8 (+0.3).
also reported in GPT-6 Astra System Card - HealthBench (Deployment Safety Hub) (system card, 36.3, mirror, section 6.1, HTML rendering of the same card) - GPT-5 OpenAI0.347as of 2026-06length-adjusted, max reasoning effort (41.6 unadjusted, 2,880 mean response chars); GPT-5.6 system card Table 6, column GPT-5. The GPT-5 launch system card printed 46.2% raw for gpt-5-thinking.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
also reported in GPT-5.6 Preview System Card (system card, 34.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5); GPT-5.5 System Card (system card, 34.7, Section 5 Health, Table 7, column GPT-5); GPT-5 System Card (system card, 46.2, conflicting, p. 18, section 3.10 Health, Figure 6 (HealthBench Hard, raw score %)) - GPT-5.2 OpenAI0.343as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (38.9 unadjusted, 2585 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
also reported in GPT-5.6 Preview System Card (system card, 34.3, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.2) - GPT-5.6 Sol OpenAI0.331as of 2026-06length-adjusted, max reasoning effort (31.1 unadjusted, 1,751 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 490.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
also reported in GPT-5.6 Preview System Card (system card, 33.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL) - GPT-5.6 Terra OpenAI0.327as of 2026-06length-adjusted, max reasoning effort (34.3 unadjusted, 2,199 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
also reported in GPT-5.6 Preview System Card (system card, 32.7, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA) - GPT-5.6 Luna OpenAI0.320as of 2026-06length-adjusted, max reasoning effort (31.4 unadjusted, 1,923 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 491.GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
also reported in GPT-5.6 Preview System Card (system card, 32.0, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA) - GPT-5.5 OpenAI0.315as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (33.8 unadjusted, 2289 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
also reported in GPT-5.6 Preview System Card (system card, 31.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.5) - GPT-5.6 Sol (August) OpenAI0.314as of 2026-08ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (27.1 unadjusted, 1,450 chars)GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard 20.2 (17.8, 1,693) 22.9 (21.3, 1,794) 31.4 (27.1, 1,450) 28.7 (24.9, 1,523)
- GPT OSS 120B OpenAI0.300as of 2025-08raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
- publisher
- OpenAI
- locator
- Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-120b high
- reported by
- vendor-reported
- published
- 2025-08-05
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard 22.8 26.9 30.0 9.0 12.9 10.8
- GPT-5.4 OpenAI0.291as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (30.3 unadjusted, 2161 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.4
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
also reported in GPT-5.6 Preview System Card (system card, 29.1, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.4) - GPT-5.6 Luna (August) OpenAI0.287as of 2026-08ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (24.9 unadjusted, 1,523 chars)GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard 20.2 (17.8, 1,693) 22.9 (21.3, 1,794) 31.4 (27.1, 1,450) 28.7 (24.9, 1,523)
- GPT-5.3 Chat OpenAI0.259as of 2026-03raw score (no length adjustment), GPT-5.3 Instant system card Table 3, column GPT-5.3-INSTANT; OpenAI later cards print 20.2 length-adjusted (17.8 unadjusted) for the re-run model.GPT-5.3 Instant System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 4.1 HealthBench, Table 3: HealthBench, row Hard, column GPT-5.3-INSTANT
- reported by
- vendor-reported
- published
- 2026-03-02
- retrieved
- 2026-09-07
- confidence
- verified
Hard 26.8% 25.9%
also reported in GPT-5.5 Instant System Card (system card, 20.2, conflicting, Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.3 INSTANT); GPT-5.6 - August Updates (system card addendum) (system card, 20.2, conflicting, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.3 Instant) - GPT-5.1 OpenAI0.254as of 2026-06length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (41.4 unadjusted, 4049 chars)GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.1
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard length-adjusted 34.7 (41.6, 2880) 25.4 (41.4, 4049) 34.3 (38.9, 2585) 29.1 (30.3, 2161) 31.5 (33.8, 2289) 33.1 (31.1, 1751) 32.7 (34.3, 2199) 32.0 (31.4, 1923)
also reported in GPT-5.6 Preview System Card (system card, 25.4, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.1) - GPT-5.5 Instant OpenAI0.229as of 2026-05length-adjusted (21.3 unadjusted, 1,794 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.GPT-5.5 Instant System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT
- reported by
- vendor-reported
- published
- 2026-05-05
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard 21.6 (23.0, 2,181) 23.3 (23.5, 2,022) 20.2 (17.8, 1,693) 22.9 (21.3, 1,794)
also reported in GPT-5.6 - August Updates (system card addendum) (system card, 22.9, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant) - GPT OSS 20B OpenAI0.108as of 2025-08raw score (%), reasoning level high; gpt-oss model card Table 3. Not length-adjusted, unlike the GPT-5.x rows on this board.gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
- publisher
- OpenAI
- locator
- Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench Hard, column gpt-oss-20b high
- reported by
- vendor-reported
- published
- 2025-08-05
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench Hard 22.8 26.9 30.0 9.0 12.9 10.8
HealthBench
18 rows · 0-100 rubric-point percentage (some sites display 0-1)Board source: GPT-5.6 System Card · official page: deploymentsafety.openai.com
- Claude Opus 5 Anthropic67.1as of 2026-07raw/unadjusted, Anthropic system-card protocol (max effort, no tools, five trials); length-adjusted 57.8; mirrored on benchlm.aiSystem Card: Claude Opus 5 · system card · first-party
- publisher
- Anthropic
- locator
- p. 188, section 8.15.1 HealthBench results; Figure 8.15.1.A
- reported by
- vendor-reported
- published
- 2026-07-24
- retrieved
- 2026-09-07
- confidence
- verified
Claude Opus 5 achieved a raw score of 67.1%, which is the highest among all Claude models, ahead of Claude Mythos 5 at 62.5%, Claude Opus 4.8 at 58.8%, and Claude Sonnet 5 at 59.2%. After length adjustment, which penalizes verbose model responses, Claude Opus 5 achieved a score of 57.8%.
also reported in Claude Fable 5.1 and Claude Mythos 5.1 System Card (system card, 67.1%, p. 198, sec. 8.17.1); GPT-5.6 Preview System Card (system card, 67.1) - Baichuan-M3 Baichuan65.1as of 2026-02self-run in Baichuan-M3 paper (arXiv 2602.06570)Baichuan-M3 Technical Report · paper
- publisher
- Baichuan
- locator
- p. 23, section 4.2.1 HealthBench-Main; Table on p. 25 (Model / HealthBench Score)
- reported by
- vendor-reported
- published
- 2026-02-06
- retrieved
- 2026-09-07
- confidence
- verified
On the comprehensive HealthBench Total, Baichuan-M3 achieves a score of 65.1, surpassing the runner-up GPT-5.2-High (63.3) by a clear margin.
also reported in Baichuan-M3 Technical Report (paper, 65.1, p. 23, section 4.2.1 HealthBench-Main (Figure 7); also Table 2, p. 25 (Baichuan-M3-235B, HealthBench Score 65.1)) - GPT-5.2-High OpenAI63.3as of 2026-02raw score as run by Baichuan in the M3 technical report, not an OpenAI-reported number; OpenAI own GPT-5.2 figure is 56.8 length-adjusted (60.7 unadjusted) in the GPT-5.6 system card.Baichuan-M3 Technical Report · paper
- publisher
- Baichuan
- locator
- p. 25, HealthBench-Hallu table (Model / HealthBench Score column); also p. 23 prose
- reported by
- independent run
- published
- 2026-02-06
- retrieved
- 2026-09-07
- confidence
- verified
GPT-5.2-High 63.3 2.37% 2.78%
also reported in Baichuan-M3 Technical Report (paper, 63.3, p. 23, section 4.2.1 HealthBench-Main (Figure 7); also Table 2, p. 25 (GPT-5.2-High, HealthBench Score 63.3)); GPT-5.6 System Card (system card, 56.8, conflicting, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2) - Claude Fable 5 Anthropic62.7as of 2026-06length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (40k token budget); printed in the Mythos 5 column of Table 8.1.A (Fable 5 column '-'); raw 62.5% per Opus 5 cardClaude Fable 5 and Claude Mythos 5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- p. 252, Table 8.1.A, row HealthBench, column Mythos 5; Figure 8.18.1.A p. 297
- reported by
- vendor-reported
- published
- 2026-06-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench 62.7 - 61.1 59.3 56.5 -
also reported in System Card: Claude Opus 5 (system card, 62.7%, p. 188, section 8.15.1, Figure 8.15.1.A); Claude Fable 5.1 and Claude Mythos 5.1 System Card (system card, 61.2% (raw), conflicting, p. 198, sec. 8.17.1) - Claude Opus 4.8 Anthropic59.3as of 2026-06length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Fable/Mythos 5 card; raw 58.8% per Opus 5 card; not in the Opus 4.8 card itselfClaude Fable 5 and Claude Mythos 5 System Card · system card · first-party
- publisher
- Anthropic
- locator
- p. 252, Table 8.1.A, row HealthBench, column Opus 4.8
- reported by
- vendor-reported
- published
- 2026-06-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench 62.7 - 61.1 59.3 56.5 -
also reported in System Card: Claude Opus 5 (system card, 59.3%, p. 188, section 8.15.1, Figure 8.15.1.A) - Claude Sonnet 5 Anthropic58.7%as of 2026-06length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; figure-only in the Sonnet 5 card; raw 59.2% per Opus 5 cardSystem Card: Claude Sonnet 5 · system card · first-party
- publisher
- Anthropic
- locator
- p. 138, section 8.12.1 HealthBench results, Figure 8.12.1.A bar label (no prose or table number)
- reported by
- vendor-reported
- published
- 2026-06-30
- retrieved
- 2026-09-07
- confidence
- verified
[Figure 8.12.1.A] HealthBench length-adjusted scores. All Claude models used adaptive thinking at max effort. (bar label: Claude Sonnet 5 58.7%)
also reported in System Card: Claude Opus 5 (system card, 58.7%, p. 188, section 8.15.1, Figure 8.15.1.A) - GPT-5.3 Chat OpenAI54.1%as of 2026-03raw score (no length adjustment), GPT-5.3 Instant system card Table 3 column GPT-5.3-INSTANT; later OpenAI cards print 49.6 length-adjusted (47.9 unadjusted)GPT-5.3 Instant System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 4.1 HealthBench, Table 3: HealthBench, row HealthBench, column GPT-5.3-INSTANT
- reported by
- vendor-reported
- published
- 2026-03-02
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench 55.4% 54.1%
also reported in GPT-5.5 Instant System Card (system card, 49.6, conflicting, Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.3 INSTANT) - Claude Fable 5.1 Anthropic60%as of 2026-09length-adjusted (method published in OpenAI's GPT-5.5 System Card); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 66.7%).Claude Fable 5.1 and Claude Mythos 5.1 System Card · system card · first-party
- publisher
- Anthropic
- locator
- p. 198, sec. 8.17.1 / Figure 8.17.1.A
- reported by
- vendor-reported
- published
- 2026-09-01
- retrieved
- 2026-09-07
- confidence
- verified
On HealthBench, Claude Fable 5.1 achieved a raw score of 66.7%, ahead of Claude Fable 5 at 61.2% and Claude Sonnet 5 at 59.2%, and behind Claude Opus 5 at 67.1%. After length adjustment, which penalizes verbose model responses, Fable 5.1 achieved a score of 60%.
- GPT-6 Astra OpenAI58.1as of 2026-09length-adjusted, max reasoning effort (59.7 unadjusted, 2,258 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'.GPT-6 Astra System Card · system card · first-party
- publisher
- OpenAI
- locator
- p. 19, sec. 6.1 (Table 6 cell, column 'gpt-6 Astra': 58.1 (59.7, 2258))
- reported by
- vendor-reported
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
Astra has a length-adjusted HealthBench Professional score of 63.4 (+2.9 relative to GPT-5.6 Sol), HealthBench score of 58.1 (+1.1), HealthBench Hard score of 36.3 (+3.2), and HealthBench Consensus score of 95.8 (+0.3).
also reported in GPT-6 Astra System Card - HealthBench (Deployment Safety Hub) (system card, 58.1, mirror, section 6.1, HTML rendering of the same card) - GPT OSS 120B OpenAI57.6as of 2025-08reasoning level high, raw score (%), gpt-oss model card Table 3 (low 53.0, medium 55.9)gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
- publisher
- OpenAI
- locator
- Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench, column gpt-oss-120b high
- reported by
- vendor-reported
- published
- 2025-08-05
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench 53.0 55.9 57.6 40.4 41.8 42.5
- GPT-5.6 Sol OpenAI57.0as of 2026-06length-adjusted, max reasoning effort (55.6 unadjusted), GPT-5.6 system card 2026-07-09GPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench length-adjusted 57.7 (63.1, 2904) 50.9 (64.2, 4222) 56.8 (60.7, 2645) 54.0 (55.7, 2275) 56.5 (58.4, 2313) 57.0 (55.6, 1764) 57.0 (58.7, 2285) 55.8 (55.4, 1930)
also reported in GPT-5.6 Preview System Card (system card, 57.0, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL) - GPT-5.6 Terra OpenAI57.0as of 2026-06length-adjusted (58.7 unadjusted), max reasoning effortGPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench length-adjusted 57.7 (63.1, 2904) 50.9 (64.2, 4222) 56.8 (60.7, 2645) 54.0 (55.7, 2275) 56.5 (58.4, 2313) 57.0 (55.6, 1764) 57.0 (58.7, 2285) 55.8 (55.4, 1930)
also reported in GPT-5.6 Preview System Card (system card, 57.0, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA) - GPT-5.5 OpenAI56.5as of 2026-04length-adjusted (58.4 unadjusted), comparison row in GPT-5.6 system cardGPT-5.5 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5 Health, Table 7 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
- reported by
- vendor-reported
- published
- 2026-04-23
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench length-adjusted 57.7 (63.1, 2904) 50.9 (64.2, 4222) 56.8 (60.7, 2645) 54.0 (55.7, 2275) 56.5 (58.4, 2313)
also reported in GPT-5.6 Preview System Card (system card, 56.5); GPT-5.6 System Card (system card, 56.5, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5) - GPT-5.6 Luna OpenAI55.8as of 2026-06length-adjusted (55.4 unadjusted), max reasoning effortGPT-5.6 System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench length-adjusted 57.7 (63.1, 2904) 50.9 (64.2, 4222) 56.8 (60.7, 2645) 54.0 (55.7, 2275) 56.5 (58.4, 2313) 57.0 (55.6, 1764) 57.0 (58.7, 2285) 55.8 (55.4, 1930)
also reported in GPT-5.6 Preview System Card (system card, 55.8, Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA) - GPT-5.6 Sol (August) OpenAI55.0as of 2026-08ChatGPT production/Instant deployment setting, length-adjusted (52.1 unadjusted), GPT-5.6 August Updates PDF 2026-08-06GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench 49.6 (47.9, 1,724) 51.4 (50.9, 1,922) 55.0 (52.1, 1,514) 53.3 (50.7, 1,567)
also reported in GPT-5.6 Preview System Card (system card, 55.0) - GPT-5.6 Luna (August) OpenAI53.3as of 2026-08ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (50.7 unadjusted, 1,567 chars)GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench 49.6 (47.9, 1,724) 51.4 (50.9, 1,922) 55.0 (52.1, 1,514) 53.3 (50.7, 1,567)
- GPT-5.5 Instant OpenAI51.4as of 2026-05length-adjusted, GPT-5.5 Instant system card Table 5 column GPT-5.5 INSTANT (50.9 unadjusted, 1,922 chars); same number in GPT-5.6 August Updates p. 11GPT-5.5 Instant System Card · system card · first-party
- publisher
- OpenAI
- locator
- Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT
- reported by
- vendor-reported
- published
- 2026-05-05
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench 49.6 (50.8, 2,208) 50.6 (51.5, 2,145) 49.6 (47.9, 1,724) 51.4 (50.9, 1,922)
also reported in GPT-5.6 - August Updates (system card addendum) (system card, 51.4, p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant) - GPT OSS 20B OpenAI42.5as of 2025-08reasoning level high, raw score (%), gpt-oss model card Table 3 (low 40.4, medium 41.8)gpt-oss-120b & gpt-oss-20b Model Card · model card · first-party
- publisher
- OpenAI
- locator
- Section 2, Table 3: Evaluations across multiple benchmarks and reasoning levels, row HealthBench, column gpt-oss-20b high
- reported by
- vendor-reported
- published
- 2025-08-05
- retrieved
- 2026-09-07
- confidence
- verified
HealthBench 53.0 55.9 57.6 40.4 41.8 42.5
Health Optimization Bench
16 rows · 0-100 rubric creditBoard source: healthoptimizationbench.com · full board: healthoptimizationbench.com
- Claude Fable 5.1 Anthropic84.9as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 84.9, 95% CI 80.1 to 89.3, n 89 (values as published in the site snapshot of 2026-09-04)
- Claude Fable 5 Anthropic83.8as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 83.8, 95% CI 79.1 to 88.1, n 89 (values as published in the site snapshot of 2026-09-04)
- Grok 4.6 xAI81.3as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 81.3, 95% CI 77 to 85.3, n 89 (values as published in the site snapshot of 2026-09-04)
- Claude Opus 5 Anthropic78.3as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 78.3, 95% CI 72.9 to 83.4, n 89 (values as published in the site snapshot of 2026-09-04)
- GPT-5.6 Sol (max) OpenAI77.9as of 2026-09max reasoning effortHealth Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 77.9, 95% CI 73.3 to 82.2, n 89 (values as published in the site snapshot of 2026-09-04)
- Kimi K3 Moonshot AI77.8as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 77.8, 95% CI 73 to 82.4, n 89 (values as published in the site snapshot of 2026-09-04)
- GPT-5.6 Sol (high) OpenAI77.7as of 2026-09high reasoning effortHealth Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 77.7, 95% CI 73.1 to 82, n 89 (values as published in the site snapshot of 2026-09-04)
- Muse Spark Meta68.4as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 68.4, 95% CI 62.8 to 73.9, n 89 (values as published in the site snapshot of 2026-09-04)
- Gemini 3.6 Google57.9as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 57.9, 95% CI 51.2 to 64.3, n 89 (values as published in the site snapshot of 2026-09-04)
- Inkling Thinking Machines53.9as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 53.9, 95% CI 47.7 to 60.1, n 89 (values as published in the site snapshot of 2026-09-04)
- Claude Sonnet 5 Anthropic50.1as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 50.1, 95% CI 44.3 to 56.1, n 89 (values as published in the site snapshot of 2026-09-04)
- MiniMax M3 MiniMax35.5as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 35.5, 95% CI 29.6 to 41.6, n 89 (values as published in the site snapshot of 2026-09-04)
- MAI Thinking Microsoft AI33.0as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 33.0, 95% CI 26.8 to 39.3, n 89 (values as published in the site snapshot of 2026-09-04)
- GLM 5.2 Zhipu29.2as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 29.2, 95% CI 23.9 to 34.9, n 89 (values as published in the site snapshot of 2026-09-04)
- Mistral Medium 3.5 Mistral14.4as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 14.4, 95% CI 10.6 to 18.5, n 89 (values as published in the site snapshot of 2026-09-04)
- Nemotron 3.5 Lightning NVIDIA7.8as of 2026-09Health Optimization Bench: Arcophos harness run · Arcophos run · first-party
- publisher
- Arcophos
- locator
- tasksets/glp1-ev/analysis.json (regenerate with benchdb import-hob on the machine holding the harness tasksets; this row was verified against the committed site snapshot), snapshot 2026-09-04
- reported by
- Arcophos run
- retrieved
- not recorded
- confidence
- verified
score 7.8, 95% CI 4.7 to 11.3, n 89 (values as published in the site snapshot of 2026-09-04)
MAST (Medical AI Superintelligence Test)
8 rows · percentage compositeBoard source: MAST: Medical AI Superintelligence Test leaderboard (General board)
- GPT-5.6 Sol OpenAI60.2%as of 2026-08MAST in preview; 'exact scores may change'MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
1 GPT-5.6 Sol OpenAI 60.2%
- Kimi K3 Moonshot AI60.1%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
2 Kimi K3 Moonshot AI 60.1%
- Gemini 3.6 Flash Google59.3%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
3 Gemini 3.6 Flash Google 59.3%
- Gemini 3.1 Pro Google58.9%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
4 Gemini 3.1 Pro Google 58.9%
- Qwen3.5 397B A17B Alibaba57.9%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
5 Qwen3.5 397B A17B Alibaba 57.9%
- Claude Opus 5 Anthropic57.1%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
6 Claude Opus 5 Anthropic 57.1%
- Claude Sonnet 5 Anthropic56.6%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
7 Claude Sonnet 5 Anthropic 56.6%
- Grok 4.3 xAI53.7%as of 2026-08MAST: Medical AI Superintelligence Test leaderboard (General board) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
8 Grok 4.3 xAI 53.7%
MedHELM
10 rows · mean win rate 0-1Board source: MedHELM leaderboard (medhelm.org), v5.0.0
- Gemini 3.1 Pro (Preview) Google0.652as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-07
- confidence
- verified
1 Gemini 3.1 Pro (Preview) Google 0.652
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.6520833333333333, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - Gemini 3.5 Flash Google0.642as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-07
- confidence
- verified
2 Gemini-3.5-flash Google 0.642
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.6416666666666667, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - Muse Spark (2026-04-08) Meta0.621as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-07
- confidence
- verified
3 Muse Spark (2026-04-08) Meta 0.621
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.6208333333333333, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - GPT-5.4 mini OpenAI0.552as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-07
- confidence
- verified
4 GPT-5.4 mini OpenAI 0.552
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.5520833333333334, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - GPT-5.4 (2026-03-05) OpenAI0.538as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-07
- confidence
- verified
5 GPT-5.4 (2026-03-05) OpenAI 0.538
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.5375, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - Gemini 2.5 Pro Google0.529as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-07
- confidence
- verified
6 Gemini 2.5 Pro (05-06 preview) Google 0.529
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.5291666666666667, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - DeepSeek R1 DeepSeek0.485as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-07
- confidence
- verified
7 DeepSeek R1 DeepSeek 0.485
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.48541666666666666, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - Claude 4.6 Opus Anthropic0.456as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-07
- confidence
- verified
8 Claude 4.6 Opus Anthropic 0.456
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.45625, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - Claude 3.7 Sonnet Anthropic0.45as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-07
- confidence
- verified
9 Claude 3.7 Sonnet (20250219) Anthropic 0.45
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.45, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate') - Gemini 2.0 Flash Google0.342as of 2026-05MedHELM leaderboard (medhelm.org), v5.0.0 · official leaderboard · first-party
- publisher
- Stanford CRFM (MedHELM)
- locator
- medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
- reported by
- official leaderboard
- published
- 2026-05-14
- retrieved
- 2026-09-07
- confidence
- verified
10 Gemini 2.0 Flash Google 0.342
also reported in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate) (official leaderboard, 0.3416666666666667, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate')
First, Do NOHARM (v2)
12 rows · percentage safety scoreBoard source: MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · paper: arxiv.org
- Muse Spark 1.1 Meta79.7%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
5Muse Spark 1.1Meta 79.7%
- Claude Opus 5 Anthropic74.6%as of 2026-08v2 run on ARISE; 19 models on the boardMAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
6Claude Opus 5Anthropic 74.6%
- Kimi K3 Moonshot AI74.0%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
7Kimi K3OSSMoonshot AI 74.0%
- GPT-5.6 Sol OpenAI70.1%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
8GPT-5.6 SolOpenAI 70.1%
- GPT-5.5 OpenAI70.0%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
9GPT-5.5OpenAI 70.0%
- GPT-5 OpenAI68.6%as of 2026-08from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI); not in the Latest Flagships rankingMAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, Model Leaderboard (Top 10 shown), row 10, SAFETY column
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
10 GPT-5 72.6%±2.3 68.6%±4.5 80.3%±4.4 30.1%±9.0 44.2%±2.5 45.4%±1.3 73.1%±3.9 73.7%±2.7 46.6%±0.0
- Claude Fable 5 Anthropic65.0%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
10Claude Fable 5Anthropic 65.0%
- Gemini 3.1 Pro Google62.6%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
11Gemini 3.1 ProGoogle 62.6%
- Gemini 2.5 Pro Google61.9%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
12Gemini 2.5 ProGoogle 61.9%
- Qwen3.5 397B A17B Alibaba61.1%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
13Qwen3.5 397B A17BOSSAlibaba 61.1%
- Kimi K2.6 Moonshot AI59.1%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
14Kimi K2.6OSSMoonshot AI 59.1%
- DeepSeek R1 DeepSeek55.8%as of 2026-08MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) · official leaderboard · first-party
- publisher
- ARISE AI Research Network
- locator
- arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
- reported by
- official leaderboard
- published
- 2026-08-15
- retrieved
- 2026-09-07
- confidence
- verified
16DeepSeek R1OSSDeepSeek 55.8%
HealthAgentBench
11 rows · mean task success rateBoard source: HealthAgentBench leaderboard · paper: arxiv.org
- Claude Code (Opus 5) Anthropic55%as of 2026-07$3.3/task; harness+model evaluated jointlyHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 1, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-07
- confidence
- verified
1 Claude Code (Opus 5) 55% $3.3
also reported in HealthAgentBench detailed results (official leaderboard, 55%, Detailed results table, rank 1) - Codex (GPT-5.6-sol) OpenAI45%as of 2026-07$5.2/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 2, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-07
- confidence
- verified
2 Codex (GPT-5.6-sol) 45% $5.2
also reported in HealthAgentBench detailed results (official leaderboard, 45%, Detailed results table, rank 2) - Codex (GPT 5.5) OpenAI42%as of 2026-07$2.8/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 3, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-07
- confidence
- verified
3 Codex (GPT 5.5) 42% $2.8
also reported in HealthAgentBench detailed results (official leaderboard, 42%, Detailed results table, rank 3); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 42%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Copilot (Opus 4.8) Microsoft/Anthropic36%as of 2026-07$3.1/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 4, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-07
- confidence
- verified
4 Copilot (Opus 4.8) 36% $3.1
also reported in HealthAgentBench detailed results (official leaderboard, 36%, Detailed results table, rank 4); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 36%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Copilot (GPT 5.5) Microsoft/OpenAI35%as of 2026-07$2.6/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 5, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-07
- confidence
- verified
5 Copilot (GPT 5.5) 35% $2.6
also reported in HealthAgentBench detailed results (official leaderboard, 35%, Detailed results table, rank 5); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 35%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Claude Code (Opus 4.8) Anthropic32%as of 2026-07$4.0/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 6, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-07
- confidence
- verified
6 Claude Code (Opus 4.8) 32% $4.0
also reported in HealthAgentBench detailed results (official leaderboard, 32%, Detailed results table, rank 6); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 32%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Codex (GPT 5.4) OpenAI28%as of 2026-07$1.3/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 7, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-07
- confidence
- verified
7 Codex (GPT 5.4) 28% $1.3
also reported in HealthAgentBench detailed results (official leaderboard, 28%, Detailed results table, rank 7); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 28%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Claude Code (Opus 4.7) Anthropic27%as of 2026-07$4.8/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 8, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-07
- confidence
- verified
8 Claude Code (Opus 4.7) 27% $4.8
also reported in HealthAgentBench detailed results (official leaderboard, 27%, Detailed results table, rank 8); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 27%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Claude Code (Opus 4.6) Anthropic19%as of 2026-07$4.1/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 10, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-07
- confidence
- verified
10 Claude Code (Opus 4.6) 19% $4.1
also reported in HealthAgentBench detailed results (official leaderboard, 19%, Detailed results table, rank 10); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 19%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Claude Code (Sonnet 4.6) Anthropic17%as of 2026-07$2.9/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 11, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-07
- confidence
- verified
11 Claude Code (Sonnet 4.6) 17% $2.9
also reported in HealthAgentBench detailed results (official leaderboard, 17%, Detailed results table, rank 11); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 17%, p. 11, Figure 4 (pooled task success rate, ten agents)) - Codex (GPT 5.4 Mini) OpenAI16%as of 2026-07$0.6/taskHealthAgentBench leaderboard · official leaderboard · first-party
- publisher
- Microsoft Research (HealthAgentBench)
- locator
- Leaderboard table (homepage), rank 12, Success Rate column
- reported by
- official leaderboard
- published
- 2026-07-27
- retrieved
- 2026-09-07
- confidence
- verified
12 Codex (GPT 5.4 Mini) 16% $0.6
also reported in HealthAgentBench detailed results (official leaderboard, 16%, Detailed results table, rank 12); HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (arXiv 2606.31179v1) (paper, 16%, p. 11, Figure 4 (pooled task success rate, ten agents))
CHI-Bench
12 rows · pass@1 with binary 0/1 rewardBoard source: CHI-Bench leaderboard (actAVA)
- erius + claude-opus-5 Humana (harness) / Anthropic (model)54.7%as of 2026-08community-submitted harness config validated by automated workspace judgeCHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- Leaderboard table, All Domains, rank 01, Accuracy column
- reported by
- independent run
- published
- 2026-08-12
- retrieved
- 2026-09-07
- confidence
- verified
01 | erius | submitted by Michael Johnson (MJ) | claude-opus-5 | Proprietary | 54.7% | 72.0% | 36.0% | 56.0% | 2026-07-26
- claude-code + claude-opus-5 Anthropic37.3%as of 2026-08CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- Leaderboard table, All Domains, rank 03, Accuracy column
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-07
- confidence
- verified
03 | claude-code | claude-opus-5 | Proprietary | 37.3% | 20.0% | 32.0% | 60.0% | 2026-07-24
- erius + claude-opus-4-8 Humana / Anthropic37.3%as of 2026-08CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- Leaderboard table, All Domains, rank 02, Accuracy column
- reported by
- independent run
- published
- 2026-08-12
- retrieved
- 2026-09-07
- confidence
- verified
02 | erius | submitted by Michael Johnson (MJ) | claude-opus-4-8 | Proprietary | 37.3% | 40.0% | 16.0% | 56.0% | 2026-06-05
- claude-code + claude-opus-4-8 Anthropic33.3%as of 2026-08CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- Leaderboard table, All Domains, rank 04, Accuracy column
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-07
- confidence
- verified
04 | claude-code | claude-opus-4-8 | Proprietary | 33.3% | 32.0% | 28.0% | 40.0% | 2026-05-28
- claude-code + claude-opus-4-6 Anthropic28.0%as of 2026-08CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- Leaderboard table, All Domains, rank 05, Accuracy column
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-07
- confidence
- verified
05 | claude-code | claude-opus-4-6 | Proprietary | 28.0% | 20.0% | 36.0% | 28.0% | 2026-05-01
- claude-code + claude-sonnet-4-6 Anthropic26.2%as of 2026-08CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- Leaderboard table, All Domains, rank 06, Accuracy column
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-07
- confidence
- verified
06 | claude-code | claude-sonnet-4-6 | Proprietary | 26.2% | 24.0% | 34.7% | 20.0% | 2026-05-01
- codex + gpt-5.6-sol OpenAI25.3%as of 2026-08CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- Leaderboard table, All Domains, rank 07, Accuracy column
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-07
- confidence
- verified
07 | codex | gpt-5.6-sol | Proprietary | 25.3% | 36.0% | 28.0% | 12.0% | 2026-07-24
- openai-agents + kimi-k3 Moonshot AI25.3%as of 2026-08CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- Leaderboard table, All Domains, rank 08, Accuracy column
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-07
- confidence
- verified
08 | openai-agents | kimi-k3 | Open-source | 25.3% | 28.0% | 32.0% | 16.0% | 2026-07-24
- claude-code + claude-opus-4-7 Anthropic24.4%as of 2026-08CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- Leaderboard table, All Domains, rank 09, Accuracy column
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-07
- confidence
- verified
09 | claude-code | claude-opus-4-7 | Proprietary | 24.4% | 24.0% | 17.3% | 32.0% | 2026-05-01
- claude-code + claude-fable-5 Anthropic24.0%as of 2026-08CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- Leaderboard table, All Domains, rank 10, Accuracy column
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-07
- confidence
- verified
10 | claude-code | claude-fable-5 | Proprietary | 24.0% | 24.0% | 24.0% | 24.0% | 2026-07-22
- codex + gpt-5.5 OpenAI20.9%as of 2026-08CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- Leaderboard table, All Domains, rank 12, Accuracy column
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-07
- confidence
- verified
12 | codex | gpt-5.5 | Proprietary | 20.9% | 29.3% | 32.0% | 1.3% | 2026-05-01
- claude-code + claude-sonnet-5 Anthropic20.0%as of 2026-08CHI-Bench leaderboard (actAVA) · official leaderboard · first-party
- publisher
- actAVA
- locator
- Leaderboard table, All Domains, rank 13, Accuracy column
- reported by
- official leaderboard
- published
- 2026-08-12
- retrieved
- 2026-09-07
- confidence
- verified
13 | claude-code | claude-sonnet-5 | Proprietary | 20.0% | 24.0% | 24.0% | 12.0% | 2026-07-06
MedCode (Vals AI)
11 rows · percentage accuracy 0-100Board source: Vals AI MedCode leaderboard
- Claude Opus 5 Anthropic63.57%as of 2026-09Vals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (Overall), rank 1, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
1 Claude Opus 5 63.57% ±1.99 $5/$25 24.50s
- Gemini 3.1 Pro Preview (02/26) Google59.06%as of 2026-09Vals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (Overall), rank 2, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
2 Gemini 3.1 Pro Preview (02/26) 59.06% ±2.00 $2/$12 38.52s
- Claude Fable 5 Anthropic56.07%as of 2026-09Vals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (Overall), rank 3, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
3 Claude Fable 5 56.07% ±2.20 $10/$50 91.44s
- Gemini 3 Flash (12/25) Google55.92%as of 2026-09Vals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (Overall), rank 4, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
4 Gemini 3 Flash (12/25) 55.92% ±2.11 $0.5/$3 44.15s
- Gemini 3.5 Flash Google55.83%as of 2026-09Vals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (Overall), rank 5, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
5 Gemini 3.5 Flash 55.83% ±2.11 $1.5/$9 25.29s
- Claude Opus 4.7 Anthropic54.86%as of 2026-09Vals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (Overall), rank 6, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
6 Claude Opus 4.7 54.86% ±2.21 $5/$25 54.25s
- Claude Fable 5.1 Anthropic53.51%as of 2026-09Vals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (Overall), rank 7, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
7 Claude Fable 5.1 53.51% ±2.17 $10/$50 3m35s
- Claude Opus 4.8 Anthropic53.22%as of 2026-09Vals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (Overall), rank 9, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
9 Claude Opus 4.8 53.22% ±2.17 $5/$25 105.92s
- Gemini 3.6 Flash Google53.15%as of 2026-09Vals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (Overall), rank 10, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
10 Gemini 3.6 Flash 53.15% ±2.16 $1.5/$7.5 16.68s
- GPT 5.1 OpenAI52.73%as of 2026-09Vals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (Overall), rank 11, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
11 GPT 5.1 52.73% ±2.15 $1.25/$10 54.55s
- GPT-6 Astra OpenAI48.49%as of 2026-09Vals AI MedCode leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (Overall), rank 22 of 90, Accuracy column; board 'Updated 9/3/2026'
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
"openai/gpt-6-astra":[0,{"accuracy":[0,48.486],"latency":[0,102.865],"stderr":[0,2.131],"cost_per_test":[0,0.451358],"temperature":[0,null],"top_p":[0,null],"max_output_tokens":[0,128000]
MedScribe (Vals AI)
12 rows · percentage accuracy 0-100Board source: Vals AI MedScribe leaderboard
- Claude Fable 5.1 Anthropic91.29%as of 2026-09Vals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (View: All Models, Task: Overall), rank 1, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
1 Claude Fable 5.1 91.29% ±1.95 $10/$50 3m08s
- Claude Opus 5 Anthropic90.98%as of 2026-09Vals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (View: All Models, Task: Overall), rank 2, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
2 Claude Opus 5 90.98% ±1.92 $5/$25 76.56s
- Muse Spark 1.2 Meta90.06%as of 2026-09Vals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (View: All Models, Task: Overall), rank 3, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
3 Muse Spark 1.2 90.06% ±1.96 $1.25/$4.25 61.46s
- Muse Spark 1.1 Meta88.89%as of 2026-09Vals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (View: All Models, Task: Overall), rank 5, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
5 Muse Spark 1.1 88.89% ±1.95 $1.25/$4.25 63.34s
- Claude Fable 5 Anthropic88.52%as of 2026-09Vals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (View: All Models, Task: Overall), rank 7, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
7 Claude Fable 5 88.52% ±1.95 $10/$50 119.47s
- GPT 5.1 OpenAI88.09%as of 2026-09Vals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (View: All Models, Task: Overall), rank 8, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
8 GPT 5.1 88.09% ±1.94 $1.25/$10 77.98s
- Kimi K3 Moonshot AI87.96%as of 2026-09Vals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (View: All Models, Task: Overall), rank 9, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
9 Kimi K3 87.96% ±1.89 $3/$15 2m17s
- GPT-6 Astra OpenAI87.91%as of 2026-09Vals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (View: All Models, Task: Overall), rank 10 of 92, Accuracy column; board 'Updated 9/3/2026'
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
"openai/gpt-6-astra":[0,{"accuracy":[0,87.908],"latency":[0,153.728],"stderr":[0,1.938],"cost_per_test":[0,0.581991],"temperature":[0,null],"top_p":[0,null],"max_output_tokens":[0,128000]
- MiniMax-M3 MiniMax87.25%as of 2026-09Vals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (View: All Models, Task: Overall), rank 11, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
11 MiniMax-M3 87.25% ±1.96 $0.6/$2.4 2m04s
- GPT 5.5 OpenAI86.87%as of 2026-09Vals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (View: All Models, Task: Overall), rank 13, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
13 GPT 5.5 86.87% ±1.93 $5/$30 2m13s
- Claude Opus 4.6 (Nonthinking) Anthropic86.74%as of 2026-09Vals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (View: All Models, Task: Overall), rank 14, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
14 Claude Opus 4.6 (Nonthinking) 86.74% ±1.94 $5/$25 54.32s
- Grok 4.6 xAI86.53%as of 2026-09Vals AI MedScribe leaderboard · official leaderboard · first-party
- publisher
- Vals AI
- locator
- Leaderboard table (View: All Models, Task: Overall), rank 15, Accuracy column
- reported by
- official leaderboard
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
15 Grok 4.6 86.53% ±1.96 $2/$6 70.94s
MedXpertQA (MM)
15 rows · percentage accuracy 0-100Board source: Introducing Muse Spark: Scaling Towards Personal Superintelligence · official page: benchlm.ai
- GPT-5.6 Sol OpenAI81.5Qwen-run comparison in the Qwen3.8-Max launch postQwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
- publisher
- Alibaba
- locator
- Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
- reported by
- independent run
- published
- 2026-08-02
- retrieved
- 2026-09-07
- confidence
- verified
| MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
- Gemini 3.1 Pro Google81.3%as of 2026-04Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence · launch post · first-party
- publisher
- Meta
- locator
- Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Gemini 3.1 Pro High; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
- reported by
- independent run
- published
- 2026-04-08
- retrieved
- 2026-09-07
- confidence
- verified
MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
also reported in Muse Spark Eval Methodology (model card, 81.3, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column Gemini 3.1 Pro High); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 81.3%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 1; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - Qwen3.8 Max Alibaba80.4%as of 2026-08Alibaba's own Qwen3.8 launch table; protocol not statedQwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
- publisher
- Alibaba
- locator
- Multimodal Benchmarks table, Multimodal Reasoning section, row MedXpertQA-MM, column Qwen3.8-Max; header row: | | Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus | Qwen3.8-Max |
- reported by
- vendor-reported
- published
- 2026-08-02
- retrieved
- 2026-09-07
- confidence
- verified
| MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
also reported in MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 80.4%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 2; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - Claude Fable 5 Anthropic80.0Qwen-run comparison in the Qwen3.8-Max launch postQwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
- publisher
- Alibaba
- locator
- Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
- reported by
- independent run
- published
- 2026-08-02
- retrieved
- 2026-09-07
- confidence
- verified
| MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
- Muse Spark Meta78.4%as of 2026-04Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence · launch post · first-party
- publisher
- Meta
- locator
- Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Muse Spark Thinking; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
- reported by
- vendor-reported
- published
- 2026-04-08
- retrieved
- 2026-09-07
- confidence
- verified
MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
also reported in Muse Spark Eval Methodology (model card, 78.4, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column Muse Spark Thinking); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 78.4%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 3; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - GPT-5.4 OpenAI77.1%as of 2026-04Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence · launch post · first-party
- publisher
- Meta
- locator
- Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column GPT 5.4 Xhigh; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
- reported by
- independent run
- published
- 2026-04-08
- retrieved
- 2026-09-07
- confidence
- verified
MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
also reported in Muse Spark Eval Methodology (model card, 77.1, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column GPT 5.4 Xhigh); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 77.1%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 4; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - GPT-5.2 OpenAI73.3Qwen-run comparison in the Qwen3.5-397B-A17B model cardQwen/Qwen3.5-397B-A17B model card · model card · first-party
- publisher
- Hugging Face
- locator
- Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
- reported by
- independent run
- published
- 2026-02-16
- retrieved
- 2026-09-07
- confidence
- verified
MedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
- Claude Opus 4.8 Anthropic71.7Qwen-run comparison in the Qwen3.8-Max launch postQwen3.8-Max: A New Bar for Coding and Cowork · launch post · first-party
- publisher
- Alibaba
- locator
- Full Benchmark Table, second (multimodal) table, row MedXpertQA-MM; columns Opus4.8 / Fable5 / Gemini3.1-Pro / GPT5.6-Sol / Qwen3.7-Plus / Qwen3.8-Max
- reported by
- independent run
- published
- 2026-08-02
- retrieved
- 2026-09-07
- confidence
- verified
| MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
- Qwen3.7 Plus Alibaba71.0%as of 2026-05Alibaba's own Qwen3.7 Plus launch table; protocol not statedQwen3.7-Plus: Multimodal Agent Intelligence · launch post · first-party
- publisher
- Alibaba
- locator
- Multimodal Benchmarks table, Multimodal Reasoning section, row MedXpertQA-MM, column Qwen3.7-Plus; header row: | | GPT-5.4 (xhigh) | Opus-4.6 Max | Gemini-3.1 Pro | Qwen3.6-Plus | Qwen3.7-Plus |
- reported by
- vendor-reported
- published
- 2026-05-31
- retrieved
- 2026-09-07
- confidence
- verified
| MedXpertQA-MM | 77.3 | 64.4 | 80.7 | 68.7 | 71.0 |
also reported in Qwen3.8-Max: A New Bar for Coding and Cowork (launch post, 71.0, Qwen3.8 launch post, Multimodal Benchmarks table, row MedXpertQA-MM, column Qwen3.7-Plus (same value carried forward)); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 71.0%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 5; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - Qwen3.5 397B A17B Alibaba70.0self-reported in the Qwen3.5-397B-A17B model cardQwen/Qwen3.5-397B-A17B model card · model card · first-party
- publisher
- Hugging Face
- locator
- Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
- reported by
- vendor-reported
- published
- 2026-02-16
- retrieved
- 2026-09-07
- confidence
- verified
MedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
- Qwen3.6 Plus Alibaba68.7Qwen-run comparison in the Qwen3.7-Plus launch postQwen3.7-Plus: Multimodal Agent Intelligence · launch post · first-party
- publisher
- Alibaba
- locator
- Multimodal Benchmarks table, row MedXpertQA-MM; columns GPT-5.4 (xhigh) / Opus-4.6 Max / Gemini-3.1 Pro / Qwen3.6-Plus / Qwen3.7-Plus
- reported by
- vendor-reported
- published
- 2026-05-31
- retrieved
- 2026-09-07
- confidence
- verified
| MedXpertQA-MM | 77.3 | 64.4 | 80.7 | 68.7 | 71.0 |
- Grok 4.20 xAI65.8%as of 2026-04Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence · launch post · first-party
- publisher
- Meta
- locator
- Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Grok 4.2 Reasoning; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
- reported by
- independent run
- published
- 2026-04-08
- retrieved
- 2026-09-07
- confidence
- verified
MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
also reported in Muse Spark Eval Methodology (model card, 65.8, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column Grok 4.2 Reasoning); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 65.8%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 6; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - Kimi K2.5 Moonshot AI65.3Qwen-run comparison in the Qwen3.5-397B-A17B model card (K2.5-1T-A32B column)Qwen/Qwen3.5-397B-A17B model card · model card · first-party
- publisher
- Hugging Face
- locator
- Benchmark Results > Vision Language > Medical VQA, row MedXpertQA-MM; columns GPT5.2 / Claude 4.5 Opus / Gemini-3 Pro / Qwen3-VL-235B-A22B / K2.5-1T-A32B / Qwen3.5-397B-A17B
- reported by
- independent run
- published
- 2026-02-16
- retrieved
- 2026-09-07
- confidence
- verified
MedXpertQA-MM 73.3 63.6 76.0 47.6 65.3 70.0
- Claude Opus 4.6 Anthropic64.8%as of 2026-04Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)Introducing Muse Spark: Scaling Towards Personal Superintelligence · launch post · first-party
- publisher
- Meta
- locator
- Launch-post benchmark table image, HEALTH section, row MedXpertQA (MM), column Opus 4.6 Max; the same table is page 5 of the Eval Methodology PDF; protocol p. 2: 'MedXpertQA Text/Multimodal: ... The multimodal variant contains 2,000 multimodal medical questions with clinical images (X-rays, histology, dermatology, etc.) and 5 answer choices (A-E). For grading, we use gpt-oss-120b to parse the predicted answer letter from free-form text.'
- reported by
- independent run
- published
- 2026-04-08
- retrieved
- 2026-09-07
- confidence
- verified
MedXpertQA (MM) | Medical Multiple Choice | 78.4 | 64.8 | 81.3 | 77.1 | 65.8 (image table row; columns Muse Spark Thinking, Opus 4.6 Max, Gemini 3.1 Pro High, GPT 5.4 Xhigh, Grok 4.2 Reasoning)
also reported in Muse Spark Eval Methodology (model card, 64.8, p. 5, benchmark table image, HEALTH section, row MedXpertQA (MM), column Opus 4.6 Max); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 64.8%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 7; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology) - Gemma 4 12B Google48.7%as of 2026-04Google's Gemma 4 model card, Unified 12B; protocol not statedGemma 4 model card · model card · first-party
- publisher
- locator
- Benchmark Results table, Vision section, row MedXPertQA MM, column Gemma 4 12B Unified; header row: | | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) |
- reported by
- vendor-reported
- published
- 2026-04-02
- retrieved
- 2026-09-07
- confidence
- verified
| MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | \- |
also reported in Gemma 4 Technical Report (paper, 48.7, Gemma 4 Technical Report p. 6, Table 6 (vision benchmarks, thinking), row MedXPertQA MM, column 12B); MedXpertQA (MM) Leaderboard & Scores - September 2026 | BenchLM.ai (mirror, 48.7%, mirror, BENCHMARK SCORE TABLE (8 MODELS), rank 8; the page says 'We mirror the published score view for MedXpertQA (MM)' and 'Updated September 4, 2026', citing Meta's Muse Spark Eval Methodology)
Artificial Analysis Healthcare & Medical Index
12 rows · index scoreBoard source: Best AI for Healthcare & Medical: LLM Leaderboard
- Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) Anthropic56as of 2026-09Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- embedded JSON-LD FAQPage, answer to "Which AI is best for doctors and clinicians?"; the Score chart JSON-LD carries the same entry as {"label":"Claude Fable 5.1 (max with fallback)","score":56.35682697075397,"detailsUrl":"/models/claude-fable-5-1"}; chart header "23 of 72 models"; retrieved 2026-09-07
- reported by
- official leaderboard
- retrieved
- 2026-09-08
- confidence
- verified
Based on the Artificial Analysis Healthcare & Medical Index, the top-performing AI models for healthcare and medical work are currently Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (56), Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (52), and Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (52). Rankings are updated as new models are released.
- Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Anthropic52as of 2026-09Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- embedded JSON-LD FAQPage, answer to "Which AI is best for doctors and clinicians?"; the Score chart JSON-LD carries the same entry as {"label":"Claude Fable 5 (with fallback)","score":52.01006901640588,"detailsUrl":"/models/claude-fable-5"}; chart header "23 of 72 models"; retrieved 2026-09-07
- reported by
- official leaderboard
- retrieved
- 2026-09-08
- confidence
- verified
Based on the Artificial Analysis Healthcare & Medical Index, the top-performing AI models for healthcare and medical work are currently Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (56), Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (52), and Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (52). Rankings are updated as new models are released.
- Claude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic52as of 2026-09all underlying benchmarks run independently by Artificial AnalysisBest AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- embedded JSON-LD FAQPage, answer to "Which AI is best for doctors and clinicians?"; the Score chart JSON-LD carries the same entry as {"label":"Claude Opus 5 (max)","score":51.66169193981395,"detailsUrl":"/models/claude-opus-5"}; chart header "23 of 72 models"; retrieved 2026-09-07
- reported by
- official leaderboard
- retrieved
- 2026-09-08
- confidence
- verified
Based on the Artificial Analysis Healthcare & Medical Index, the top-performing AI models for healthcare and medical work are currently Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (56), Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (52), and Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (52). Rankings are updated as new models are released.
- GPT-5.6 Sol (max) OpenAI44as of 2026-09Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 9 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
- reported by
- official leaderboard
- retrieved
- 2026-09-08
- confidence
- verified
{"label":"GPT-5.6 Sol (max)","score":43.7633355450504,"detailsUrl":"/models/gpt-5-6-sol"}
- Grok 4.6 (high) xAI44as of 2026-09Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 8 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
- reported by
- official leaderboard
- retrieved
- 2026-09-08
- confidence
- verified
{"label":"Grok 4.6 (high)","score":43.92324392614704,"detailsUrl":"/models/grok-4-6"}
- Kimi K3 (max) Moonshot AI44as of 2026-09Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 10 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
- reported by
- official leaderboard
- retrieved
- 2026-09-08
- confidence
- verified
{"label":"Kimi K3 (max)","score":43.62067857170797,"detailsUrl":"/models/kimi-k3"}
- GPT-5.6 Terra (max) OpenAI39as of 2026-09Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 12 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
- reported by
- official leaderboard
- retrieved
- 2026-09-08
- confidence
- verified
{"label":"GPT-5.6 Terra (max)","score":38.66673219960154,"detailsUrl":"/models/gpt-5-6-terra"}
- GPT-5.6 Luna (max) OpenAI34as of 2026-09Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 15 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
- reported by
- official leaderboard
- retrieved
- 2026-09-08
- confidence
- verified
{"label":"GPT-5.6 Luna (max)","score":34.39525059118443,"detailsUrl":"/models/gpt-5-6-luna"}
- Inkling (xhigh) Thinking Machines31as of 2026-09Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 18 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
- reported by
- official leaderboard
- retrieved
- 2026-09-08
- confidence
- verified
{"label":"Inkling","score":30.52028583441501,"detailsUrl":"/models/inkling"}
- MiniMax-M3 MiniMax31as of 2026-09Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- embedded JSON-LD Dataset "Artificial Analysis Healthcare & Medical Index" (Score chart data), entry 17 of 20 listed; chart header "23 of 72 models"; retrieved 2026-09-07
- reported by
- official leaderboard
- retrieved
- 2026-09-08
- confidence
- verified
{"label":"MiniMax-M3","score":30.552965002205703,"detailsUrl":"/models/minimax-m3"}
- Mistral Medium 3.5 Mistral17as of 2026-09Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- page data payload initialModels (RSC flight, JSON-escaped inside self.__next_f.push in the cached HTML), entry for slug "mistral-medium-3-5", rank 22 of 23; the Score chart JSON-LD lists only the top 20; chart header "23 of 72 models"; retrieved 2026-09-07
- reported by
- official leaderboard
- retrieved
- 2026-09-08
- confidence
- verified
\"slug\":\"mistral-medium-3-5\",\"name\":\"Mistral Medium 3.5\" ... \"headlineValue\":17.2489866350548
- gpt-oss-120b (high) OpenAI14as of 2026-09Best AI for Healthcare & Medical: LLM Leaderboard · official leaderboard · first-party
- publisher
- Artificial Analysis
- locator
- page data payload initialModels (RSC flight, JSON-escaped inside self.__next_f.push in the cached HTML), entry for slug "gpt-oss-120b", rank 23 of 23; the Score chart JSON-LD lists only the top 20; chart header "23 of 72 models"; retrieved 2026-09-07
- reported by
- official leaderboard
- retrieved
- 2026-09-08
- confidence
- verified
\"slug\":\"gpt-oss-120b\",\"name\":\"gpt-oss-120b (high)\" ... \"headlineValue\":14.423032488856672
PhysicianBench
9 rows · pass@1 success rate %Board source: PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · official page: arxiv.org
- GPT-5.5 OpenAI46.3 ± 1.2as of 2026-05pass@1; Pass^3 28.0PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- official leaderboard
- published
- 2026-05-04
- retrieved
- 2026-09-07
- confidence
- verified
GPT-5.5 46.3 ± 1.2 57.4 28.0 41.9
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 46.3 ± 1.2, arXiv abs page (landing page for the PDF)) - Claude Opus 4.6 Anthropic31.7 ± 2.3as of 2026-05Pass^3 18.0PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- official leaderboard
- published
- 2026-05-04
- retrieved
- 2026-09-07
- confidence
- verified
Claude Opus 4.6 31.7 ± 2.3 41.5 18.0 25.2
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 31.7 ± 2.3, arXiv abs page (landing page for the PDF)) - Claude Opus 4.7 Anthropic29.3 ± 2.5as of 2026-05Pass^3 18.0PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- official leaderboard
- published
- 2026-05-04
- retrieved
- 2026-09-07
- confidence
- verified
Claude Opus 4.7 29.3 ± 2.5 37.9 18.0 16.2
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 29.3 ± 2.5, arXiv abs page (landing page for the PDF)) - GPT-5.4 OpenAI27.7 ± 1.5as of 2026-05PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- official leaderboard
- published
- 2026-05-04
- retrieved
- 2026-09-07
- confidence
- verified
GPT-5.4 27.7 ± 1.5 37.7 13.0 39.8
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 27.7 ± 1.5, arXiv abs page (landing page for the PDF)) - Claude Sonnet 4.6 Anthropic23.0 ± 2.6as of 2026-05PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- official leaderboard
- published
- 2026-05-04
- retrieved
- 2026-09-07
- confidence
- verified
Claude Sonnet 4.6 23.0 ± 2.6 33.2 9.0 22.3
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 23.0 ± 2.6, arXiv abs page (landing page for the PDF)) - Kimi-K2.6 Moonshot AI17.0 ± 2.6as of 2026-05open sourcePhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- official leaderboard
- published
- 2026-05-04
- retrieved
- 2026-09-07
- confidence
- verified
Kimi-K2.6 17.0 ± 2.6 26.3 5.0 42.4
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 17.0 ± 2.6, arXiv abs page (landing page for the PDF)) - Qwen3.6-Plus Alibaba13.7 ± 4.0as of 2026-05PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- official leaderboard
- published
- 2026-05-04
- retrieved
- 2026-09-07
- confidence
- verified
Qwen3.6-Plus 13.7 ± 4.0 22.6 2.0 28.0
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 13.7 ± 4.0, arXiv abs page (landing page for the PDF)) - Gemini Pro 3.1 Google6.0 ± 1.0as of 2026-05PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (columns Pass@1, Pass@3, Pass^3, #Turns)
- reported by
- official leaderboard
- published
- 2026-05-04
- retrieved
- 2026-09-07
- confidence
- verified
Gemini Pro 3.1 6.0 ± 1.0 9.3 3.0 30.4
also reported in PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240) (paper, 6.0 ± 1.0, arXiv abs page (landing page for the PDF)) - Grok-4.20 xAI5.3 ± 3.2as of 2026-05PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments (arXiv 2605.02240v1 PDF) · paper
- publisher
- Stanford University (HealthRex; Liu, Chen et al.)
- locator
- p. 8, Table 2 (Proprietary Models block)
- reported by
- official leaderboard
- published
- 2026-05-04
- retrieved
- 2026-09-07
- confidence
- verified
Grok-4.20 5.3 ± 3.2 9.7 1.0 16.7
EHR-Complex
6 rows · exact-match accuracyBoard source: EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · official page: arxiv.org
- GPT-5.4 (high reasoning) OpenAI0.65as of 2026-06average over 12 intent columns; run as human-validation configuration, not in the headline 12-model tableEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 15, Table 10 (Strong commercial model results), Avg. column
- reported by
- official leaderboard
- published
- 2026-06-22
- retrieved
- 2026-09-07
- confidence
- verified
GPT-5.4 (high) 0.85 0.36 0.78 0.34 0.8 0.37 0.93 0.76 0.91 0.49 0.84 0.36 0.65
also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.65, arXiv abs page (landing page for the PDF)) - Gemini 3.1 Pro Google0.63as of 2026-06validation configurationEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 15, Table 10 (Strong commercial model results), Avg. column
- reported by
- official leaderboard
- published
- 2026-06-22
- retrieved
- 2026-09-07
- confidence
- verified
Gemini 3.1 Pro 0.87 0.32 0.67 0.38 0.84 0.37 0.92 0.63 0.92 0.44 0.83 0.34 0.63
also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.63, arXiv abs page (landing page for the PDF)) - Kimi-K2.5 Moonshot AI0.62as of 2026-06headline 12-model evaluation, top open-weightEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. column
- reported by
- official leaderboard
- published
- 2026-06-22
- retrieved
- 2026-09-07
- confidence
- verified
Kimi-K2.5 0.89 0.34 0.8 0.27 0.73 0.35 0.92 0.72 0.89 0.42 0.84 0.29 0.62
also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.62, arXiv abs page (landing page for the PDF)) - Qwen3.5-397B Alibaba0.62as of 2026-06headline evaluationEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 6, Table 3 (Evaluation Results on the EHR-Complex Test Set), Avg. column
- reported by
- official leaderboard
- published
- 2026-06-22
- retrieved
- 2026-09-07
- confidence
- verified
Qwen3.5-397B 0.88 0.39 0.72 0.3 0.8 0.38 0.91 0.63 0.91 0.45 0.78 0.32 0.62
also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.62, arXiv abs page (landing page for the PDF)) - GPT-5.4 (low reasoning) OpenAI0.58as of 2026-06validation configurationEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 15, Table 10 (Strong commercial model results), Avg. column
- reported by
- official leaderboard
- published
- 2026-06-22
- retrieved
- 2026-09-07
- confidence
- verified
GPT-5.4 (low) 0.81 0.32 0.65 0.26 0.73 0.31 0.88 0.66 0.86 0.43 0.8 0.29 0.58
also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.58, arXiv abs page (landing page for the PDF)) - Claude Sonnet 4.6 Anthropic0.36as of 2026-06validation configurationEHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301v1 PDF) · paper
- publisher
- Zhejiang University / Ant Group (Qiao, Liu, Chu et al.)
- locator
- p. 15, Table 10 (Strong commercial model results), Avg. column
- reported by
- official leaderboard
- published
- 2026-06-22
- retrieved
- 2026-09-07
- confidence
- verified
Claude Sonnet 4.6 0.59 0.21 0.45 0.12 0.45 0.12 0.83 0.2 0.49 0.33 0.4 0.11 0.36
also reported in EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (arXiv 2606.23301) (paper, 0.36, arXiv abs page (landing page for the PDF))
WHBench
6 rows · mean normalized percentage 0-100Board source: WHBench: A Women's Health Benchmark for Evaluating Frontier LLMs with Expert-in-the-Loop Validation (arXiv 2604.00024v2) · official page: arxiv.org
- Claude Opus 4.6 Anthropic72.1%as of 2026-0395% CI 69.6-74.4; evaluations run March 2026
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- official leaderboard
- published
- 2026-07-23
- retrieved
- 2026-09-07
- confidence
- verified
1 Claude Opus 4.6 72.1 [69.6, 74.4] 35.5 58.2 6.4 12.8
also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 72.1%, same paper, abstract page) - Claude Sonnet 4.6 Anthropic67.1%as of 2026-0395% CI 64.5-69.6
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- official leaderboard
- published
- 2026-07-23
- retrieved
- 2026-09-07
- confidence
- verified
2 Claude Sonnet 4.6 67.1 [64.5, 69.6] 22.7 67.4 9.9 27.0
also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 67.1%, same paper, abstract page) - GPT-5.4 OpenAI66.8%as of 2026-0395% CI 64.5-69.2
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- official leaderboard
- published
- 2026-07-23
- retrieved
- 2026-09-07
- confidence
- verified
3 GPT-5.4 66.8 [64.5, 69.2] 21.3 67.4 11.3 47.5
also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 66.8%, same paper, abstract page) - Gemini 3 Flash Preview Google64.7%as of 2026-03
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- official leaderboard
- published
- 2026-07-23
- retrieved
- 2026-09-07
- confidence
- verified
4 Gemini 3 Flash Preview 64.7 [61.7, 67.7] 25.5 62.4 12.1 32.6
also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 64.7%, same paper, abstract page) - GPT-4.1 OpenAI51.8%as of 2026-03
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- official leaderboard
- published
- 2026-07-23
- retrieved
- 2026-09-07
- confidence
- verified
11 GPT-4.1 51.8 [49.2, 54.3] 3.5 61.7 34.8 61.0
also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 51.8%, same paper, abstract page) - GPT-4o OpenAI44.6%as of 2026-03
- publisher
- Maurya, Govindgari, Kumar (Columbia University / Akhara AI)
- locator
- p. 5, Table 3 (WHBench v3.0 leaderboard)
- reported by
- official leaderboard
- published
- 2026-07-23
- retrieved
- 2026-09-07
- confidence
- verified
16 GPT-4o 44.6 [41.8, 47.4] 1.4 42.5 56.0 83.7
also reported in WHBench (arXiv 2604.00024 abstract page) (paper, 44.6%, same paper, abstract page)
HealthAdminBench
7 rows · percentage end-to-end task success 0-100Board source: HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks (arXiv 2604.09937v1) · official page: healthadminbench.stanford.edu · paper: arxiv.org
- Claude Opus 4.6 (computer-use agent) Anthropic36.3%as of 2026-04screenshot-only, detailed prompting; native CUA harness; subtask rate 78.4%
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- official leaderboard
- published
- 2026-04-10
- retrieved
- 2026-09-07
- confidence
- verified
Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate
also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 36.3%, Section 'LLMs struggle with long-horizon tasks') - GPT-5.4 (computer-use agent) OpenAI26.7%as of 2026-04screenshot-only, detailed prompting; subtask rate 82.8%
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- official leaderboard
- published
- 2026-04-10
- retrieved
- 2026-09-07
- confidence
- verified
Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate
also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 26.7%, Section 'LLMs struggle with long-horizon tasks') - Kimi K2.5 Moonshot AI15.6%as of 2026-04screenshot-only, detailed prompting
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- official leaderboard
- published
- 2026-04-10
- retrieved
- 2026-09-07
- confidence
- verified
Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate
also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 15.6%, Section 'LLMs struggle with long-horizon tasks') - Claude Opus 4.6 (standardized harness) Anthropic14.8%as of 2026-04screenshot-only, detailed prompting; authors' standardized harness, no native CUA
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- official leaderboard
- published
- 2026-04-10
- retrieved
- 2026-09-07
- confidence
- verified
Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate
- Qwen 3.5 Alibaba13.3%as of 2026-04screenshot-only, detailed prompting
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- official leaderboard
- published
- 2026-04-10
- retrieved
- 2026-09-07
- confidence
- verified
Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate
also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 13.3%, Section 'LLMs struggle with long-horizon tasks') - Gemini 3.1 Pro Google11.9%as of 2026-04screenshot-only, detailed prompting
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- official leaderboard
- published
- 2026-04-10
- retrieved
- 2026-09-07
- confidence
- verified
Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate
also reported in Introducing HealthAdminBench: AI Agents Can Diagnose Rare Diseases, But Can They Handle Your Insurance? (Kinetic Systems blog) (vendor post, 11.9%, Section 'LLMs struggle with long-horizon tasks') - GPT-5.4 (standardized harness) OpenAI5.9%as of 2026-04screenshot-only, detailed prompting; authors' standardized harness, no native CUA
- publisher
- Bedi, Welch, Steinberg et al. (Stanford University / Kinetic Systems / Stanford Health Care)
- locator
- p. 8, Figure 3(a) Task Success Rate (bar labels and values, extracted in order)
- reported by
- official leaderboard
- published
- 2026-04-10
- retrieved
- 2026-09-07
- confidence
- verified
Claude Opus 4.6 CUA GPT-5.4 CUA Kimi K2.5 Claude Opus 4.6 Qwen 3.5 Gemini 3.1 Pro GPT-5.4 36.3% 26.7% 15.6% 14.8% 13.3% 11.9% 5.9% Task Success Rate
OpenAI Dynamic Mental Health Evaluations
15 rows · compliance rate per metricBoard source: GPT-5.6 - August Updates (system card addendum)
- GPT-6 Astra OpenAI1.000as of 2026-09dynamic multi-turn adversarial user simulations, metric 'safe' (share of assistant messages that do not violate safety policies): mental health 1.000, emotional reliance 0.993, self-harm 0.989.GPT-6 Astra System Card · system card · first-party
- publisher
- OpenAI
- locator
- p. 20, Table 7 (Dynamic Benchmarks with Adversarial User Simulations), column 'gpt-6 Astra'
- reported by
- vendor-reported
- published
- 2026-09-03
- retrieved
- 2026-09-08
- confidence
- verified
Mental health 0.914 0.820 0.991 1.000
- GPT-5.5 Instant OpenAI0.999as of 2026-05mental health 0.999, emotional reliance 0.963, self-harm 0.913GPT-5.5 Instant System Card (PDF) · system card · first-party
- publisher
- OpenAI
- locator
- p. 6, Table 3, sec. 3.3
- reported by
- vendor-reported
- published
- 2026-05-05
- retrieved
- 2026-09-07
- confidence
- verified
Mental health 0.818 1.000 1.000 0.999 / Emotional reliance 0.976 0.992 0.995 0.963 / Self-harm 0.842 0.976 0.924 0.913 (gpt-5.5-instant column)
also reported in GPT-5.6 - August Updates (system card addendum) (system card, 0.996, conflicting, p. 9, sec. 3.4) - GPT-5.5 Instant (June Update) OpenAI0.991as of 2026-08mental health 0.991, emotional reliance 0.989, self-harm 0.967; measured at lowest reasoning deployment settingsGPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 9, table "Dynamic Benchmarks with Adversarial User Simulations", sec. 3.4
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-07
- confidence
- verified
Mental health 0.984 0.996 0.981 0.991 0.981 0.977 / Emotional reliance 0.987 0.965 0.988 0.989 0.961 0.965 / Self-harm 0.816 0.839 0.852 0.967 0.901 0.911 (GPT-5.5 Instant (June Update) column)
- GPT-5.6 Sol (July) OpenAI0.991as of 2026-07July system card version (API, Codex and Work deployments), dynamic multi-turn adversarial user simulations, metric not_unsafe: mental health 0.991, emotional reliance 0.953, self-harm 0.856. Distinct from the August ChatGPT production version measured in row 193 (0.981 / 0.961 / 0.901).GPT-5.6 System Card (PDF) · system card · first-party
- publisher
- OpenAI
- locator
- p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
Mental health 0.753 0.975 0.985 0.820 0.991 0.985 0.989 / Emotional reliance 0.857 0.953 0.985 0.915 0.953 0.976 0.957 / Self-harm 0.904 0.955 0.977 0.868 0.856 0.947 0.905 (gpt-5.6-sol column)
- GPT-5.6 Luna (July) OpenAI0.989as of 2026-07July system card version, dynamic multi-turn adversarial user simulations, metric not_unsafe: mental health 0.989, emotional reliance 0.957, self-harm 0.905. Distinct from the August ChatGPT production version measured in row 194 (0.977 / 0.965 / 0.911).GPT-5.6 System Card (PDF) · system card · first-party
- publisher
- OpenAI
- locator
- p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
Mental health ... 0.989 / Emotional reliance ... 0.957 / Self-harm ... 0.905 (gpt-5.6-luna column)
- GPT-5.3 Instant OpenAI0.985as of 2026-03mental health 0.985, emotional reliance 0.992, self-harm 0.911; dynamic multi-turn adversarial user simulations, metric not_unsafeGPT-5.3 Instant System Card · system card · first-party
- publisher
- OpenAI
- locator
- p. 2, Table 2
- reported by
- vendor-reported
- published
- 2026-03-03
- retrieved
- 2026-09-07
- confidence
- verified
mental health* 0.832 1.000 0.985 / emotional reliance* 0.945 0.952 0.992 / Self-harm * 0.845 0.920 0.911
also reported in GPT-5.5 Instant System Card (PDF) (system card, 1.000, conflicting, p. 6, Table 3, sec. 3.3); GPT-5.6 - August Updates (system card addendum) (system card, 0.984, conflicting, p. 9, sec. 3.4) - GPT-5.4 mini OpenAI0.985as of 2026-03mental health 0.985, emotional reliance 0.977, self-harm 0.980; dynamic multi-turn adversarial user simulations, metric not_unsafeGPT-5.4 Thinking System Card · system card · first-party
- publisher
- OpenAI
- locator
- p. 32, Table 17 (Appendix 6.1, added 2026-03-17)
- reported by
- vendor-reported
- published
- 2026-03-05
- retrieved
- 2026-09-07
- confidence
- verified
Mental health 0.753 0.975 0.985 0.985 / Emotional reliance 0.857 0.953 0.985 0.977 / Self-harm 0.904 0.955 0.977 0.980
- GPT-5.4 Thinking OpenAI0.985as of 2026-03mental health 0.985, emotional reliance 0.985, self-harm 0.977; dynamic multi-turn adversarial user simulations, metric not_unsafeGPT-5.4 Thinking System Card · system card · first-party
- publisher
- OpenAI
- locator
- p. 32, Table 17 (same values as p. 5, Table 2)
- reported by
- vendor-reported
- published
- 2026-03-05
- retrieved
- 2026-09-07
- confidence
- verified
Mental health 0.753 0.975 0.985 0.985 / Emotional reliance 0.857 0.953 0.985 0.977 / Self-harm 0.904 0.955 0.977 0.980
also reported in GPT-5.5 System Card (system card, 0.985, p. 12, Table 8 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2); GPT-5.6 System Card (PDF) (system card, 0.985, p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2) - GPT-5.6 Terra OpenAI0.985as of 2026-07mental health 0.985, emotional reliance 0.976, self-harm 0.947GPT-5.6 System Card (PDF) · system card · first-party
- publisher
- OpenAI
- locator
- p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2
- reported by
- vendor-reported
- published
- 2026-07-09
- retrieved
- 2026-09-07
- confidence
- verified
Mental health ... 0.985 / Emotional reliance ... 0.976 / Self-harm ... 0.947 (gpt-5.6-terra column)
- GPT-5.5 OpenAI0.981as of 2026-04mental health 0.981, emotional reliance 0.981, self-harm 0.937GPT-5.5 System Card · system card · first-party
- publisher
- OpenAI
- locator
- p. 12, Table 8 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2
- reported by
- vendor-reported
- published
- 2026-04-23
- retrieved
- 2026-09-07
- confidence
- verified
Mental health 0.753 0.975 0.985 0.981 / Emotional reliance 0.857 0.953 0.985 0.981 / Self-harm 0.904 0.955 0.977 0.937 (gpt-5.5 column)
also reported in GPT-5.6 System Card (PDF) (system card, 0.820, conflicting, p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2) - GPT-5.5 Instant (May Update) OpenAI0.981as of 2026-08mental health 0.981, emotional reliance 0.988, self-harm 0.852GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 9, table "Dynamic Benchmarks with Adversarial User Simulations", sec. 3.4
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-07
- confidence
- verified
Mental health 0.984 0.996 0.981 0.991 0.981 0.977 / Emotional reliance 0.987 0.965 0.988 0.989 0.961 0.965 / Self-harm 0.816 0.839 0.852 0.967 0.901 0.911 (GPT-5.5 Instant (May Update) column)
- GPT-5.6 Sol (August) OpenAI0.981as of 2026-08mental health 0.981, emotional reliance 0.961, self-harm 0.901; OpenAI flags statistically significant offline self-harm regression vs GPT-5.5 June, not reproduced onlineGPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 9, table "Dynamic Benchmarks with Adversarial User Simulations", sec. 3.4
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-07
- confidence
- verified
Mental health 0.984 0.996 0.981 0.991 0.981 0.977 / Emotional reliance 0.987 0.965 0.988 0.989 0.961 0.965 / Self-harm 0.816 0.839 0.852 0.967 0.901 0.911 (GPT-5.6 Sol (August) column)
- GPT-5.6 Luna (August) OpenAI0.977as of 2026-08mental health 0.977, emotional reliance 0.965, self-harm 0.911GPT-5.6 - August Updates (system card addendum) · system card · first-party
- publisher
- OpenAI
- locator
- p. 9, table "Dynamic Benchmarks with Adversarial User Simulations", sec. 3.4
- reported by
- vendor-reported
- published
- 2026-08-06
- retrieved
- 2026-09-07
- confidence
- verified
Mental health 0.984 0.996 0.981 0.991 0.981 0.977 / Emotional reliance 0.987 0.965 0.988 0.989 0.961 0.965 / Self-harm 0.816 0.839 0.852 0.967 0.901 0.911 (GPT-5.6 Luna (August) column)
- GPT-5.2 Thinking OpenAI0.975as of 2026-03mental health 0.975, emotional reliance 0.953, self-harm 0.955; dynamic multi-turn adversarial user simulations, metric not_unsafeGPT-5.4 Thinking System Card · system card · first-party
- publisher
- OpenAI
- locator
- p. 32, Table 17
- reported by
- vendor-reported
- published
- 2026-03-05
- retrieved
- 2026-09-07
- confidence
- verified
Mental health 0.753 0.975 0.985 0.985 / Emotional reliance 0.857 0.953 0.985 0.977 / Self-harm 0.904 0.955 0.977 0.980
also reported in GPT-5.5 System Card (system card, 0.975, p. 12, Table 8 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2); GPT-5.6 System Card (PDF) (system card, 0.975, p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2) - GPT-5.1 Thinking OpenAI0.753as of 2026-03mental health 0.753, emotional reliance 0.857, self-harm 0.904; dynamic multi-turn adversarial user simulations, metric not_unsafeGPT-5.4 Thinking System Card · system card · first-party
- publisher
- OpenAI
- locator
- p. 32, Table 17 (columns gpt-5.1-thinking, gpt-5.2-thinking, gpt-5.4-thinking, gpt-5.4-mini)
- reported by
- vendor-reported
- published
- 2026-03-05
- retrieved
- 2026-09-07
- confidence
- verified
Mental health 0.753 0.975 0.985 0.985 / Emotional reliance 0.857 0.953 0.985 0.977 / Self-harm 0.904 0.955 0.977 0.980
also reported in GPT-5.5 System Card (system card, 0.753, p. 12, Table 8 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2); GPT-5.6 System Card (PDF) (system card, 0.753, p. 16, Table 7 "Dynamic Benchmarks with Adversarial User Simulations", sec. 5.2)
Documents on file
The distinct documents the sections above draw on, with the number of index rows each one backs.
Corrections go through the same route as everything else here: a better document replaces a weaker one, the row's confidence moves, and the change is dated on the updates page. The full record, sources included, is downloadable from the data page.