Clinical Benchmarks

Claude Sonnet 5

Released 30 Jun 2026$2 input, $10 output per million tokens1.0M contextproprietaryAlso written as claude-sonnet-5Compare with other models

Clinical Benchmarks Index
61.2rank 40 of 148; 61.2 × 1 = 61.2, from 5 of 10 boards
Boards
6 of 13
Results
6
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. HealthBench Professional
    length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).
    0.578
    Rank 13 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026
  2. 56.6
    Rank 7 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
    Official leaderboard
    Measured Aug 2026

Documentation and coding

  1. MedScribe (Vals AI)
    model ID anthropic/claude-sonnet-5; temperature=1; max_output_tokens=30000; compute_effort=max
    76.05
    Rank 74 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000
    47.54
    Rank 30 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    37.4
    Rank 8 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  2. CHI-Bench
    Listed as claude-code + claude-sonnet-5
    20.0
    Rank 13 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Jul 2026

Sources

Open a line for the quote and page.

  1. 13HealthBench Professional length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Op… 0.578
    Printed as 57.8Vendor-reported, measured Jun 2026Configuration: length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).
    System Card: Claude Sonnet 5 system card, Anthropic, 30 Jun 2026. p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 5; Figure 8.12.2.A p. 139
    HealthBench Professional 57.8 44.2 51.8 -
    Every result from this document
  2. 74MedScribe (Vals AI) model ID anthropic/claude-sonnet-5; temperature=1; max_output_tokens=30000; com… 76.05
    Printed as 76.05%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-5; temperature=1; max_output_tokens=30000; compute_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 74 of 105 (Claude Sonnet 5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-5"].
    74 | Claude Sonnet 5 | 76.05%±3.05 | $2/$10 | 4m12s
    Every result from this document
  3. 30MedCode (Vals AI) model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_outp… 47.54
    Printed as 47.54%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 30 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-sonnet-5"].
    30 | Claude Sonnet 5 | 47.54%±2.27 | $2/$10 | 2m15s
    Every result from this document
  4. 7MAST (Medical AI Superintelligence Test) 56.6
    Printed as 56.6%Official leaderboard, measured Aug 2026
    MAST: Medical AI Superintelligence Test leaderboard (General board) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    7 Claude Sonnet 5 Anthropic 56.6%
    Every result from this document
  5. 8PhysicianBench Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 37.4
    Printed as 37.4%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Sonnet 5 (max).
    PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
    Every result from this document
  6. 13CHI-Bench 20.0
    Printed as 20.0%Official leaderboard, measured Jul 2026
    CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 13, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-07-06; accessed 2026-09-30. Linked submission provenance: run 2026-07-02T04:22:04Z to 2026-07-02T04:31:46Z; submitted_at 2026-07-06T05:23:12Z.
    13 | claude-code | claude-sonnet-5 | Proprietary | 20.0% | 24.0% | 24.0% | 12.0% | 2026-07-06
    Every result from this document

Other Anthropic models: Claude 3.7 Sonnet, Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.1, Claude Opus 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 4, Claude Sonnet 4.5, Claude Sonnet 4.6, Claude Sonnet 5.5