Clinical Benchmarks

Claude Sonnet 5.5

Released 28 Sep 2026$2 input, $10 output per million tokens1M contextproprietaryAlso written as claude-sonnet-5-5, anthropic/claude-sonnet-5-5Compare with other models

Clinical Benchmarks Index
91.0rank 1 of 148; 91.0 × 1 = 91.0, from 4 of 10 boards
Boards
4 of 13
Results
8
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. HealthBench Professional
    Anthropic; max effort; Opus 4.8 grader; safety classifiers enabled; length-adjusted; raw 77.1%
    0.692
    Rank 2 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Sep 2026

Documentation and coding

  1. MedScribe (Vals AI)
    model ID anthropic/claude-sonnet-5-5; temperature=1; max_output_tokens=128000; compute_effort=max
    91.10
    Rank 3 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000
    52.92
    Rank 11 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    63.2
    Rank 2 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  2. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; xhigh effort; Opus 5 rubric grader; safety classifiers enabled
    56.4
    Rank 5 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  3. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; high effort; Opus 5 rubric grader; safety classifiers enabled
    47.6
    Rank 6 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  4. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; medium effort; Opus 5 rubric grader; safety classifiers enabled
    30.0
    Rank 10 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  5. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; low effort; Opus 5 rubric grader; safety classifiers enabled
    27.2
    Rank 13 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 2HealthBench Professional Anthropic; max effort; Opus 4.8 grader; safety classifiers enabled; length-adju… 0.692
    Printed as 69.2%Vendor-reported, measured Sep 2026Configuration: Anthropic; max effort; Opus 4.8 grader; safety classifiers enabled; length-adjusted; raw 77.1%
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Section 8.15.2 and paragraph above Figure 8.15.B.
    On HealthBench Professional at max effort, length adjustment changes the ranking. After: GPT-6 Astra (70.3%) > Claude Sonnet 5.5 (69.2%) > Claude Opus 5.5 (65.6%) > Claude Fable 5.1 (62.1%).
    Every result from this document
  2. 3MedScribe (Vals AI) model ID anthropic/claude-sonnet-5-5; temperature=1; max_output_tokens=128000;… 91.10
    Printed as 91.10%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-5-5; temperature=1; max_output_tokens=128000; compute_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 3 of 105 (Claude Sonnet 5.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-5-5"].
    3 | Claude Sonnet 5.5 | 91.10%±1.96 | $2/$10 | 5m07s
    Every result from this document
  3. 11MedCode (Vals AI) model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_ou… 52.92
    Printed as 52.92%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 11 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-sonnet-5-5"].
    11 | Claude Sonnet 5.5 | 52.92%±2.12 | $2/$10 | 3m59s
    Every result from this document
  4. 2PhysicianBench Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 63.2
    Printed as 63.2%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Sonnet 5.5 (max).
    PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
    Every result from this document
  5. 5PhysicianBench Anthropic harness; 100 tasks; pass@1; xhigh effort; Opus 5 rubric grader; safet… 56.4
    Printed as 56.4%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; xhigh effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Figure 8.15.B, right panel PhysicianBench (pass@1); orange Sonnet 5.5 xhigh label (visually read printed labels).
    Sonnet 5.5 | PhysicianBench (pass@1) | xhigh 56.4%
    Every result from this document
  6. 6PhysicianBench Anthropic harness; 100 tasks; pass@1; high effort; Opus 5 rubric grader; safety… 47.6
    Printed as 47.6%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; high effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Figure 8.15.B, right panel PhysicianBench (pass@1); orange Sonnet 5.5 high label (visually read printed labels).
    Sonnet 5.5 | PhysicianBench (pass@1) | high 47.6%
    Every result from this document
  7. 10PhysicianBench Anthropic harness; 100 tasks; pass@1; medium effort; Opus 5 rubric grader; safe… 30.0
    Printed as 30.0%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; medium effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Figure 8.15.B, right panel PhysicianBench (pass@1); orange Sonnet 5.5 medium label (visually read printed labels).
    Sonnet 5.5 | PhysicianBench (pass@1) | medium 30.0%
    Every result from this document
  8. 13PhysicianBench Anthropic harness; 100 tasks; pass@1; low effort; Opus 5 rubric grader; safety… 27.2
    Printed as 27.2%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; low effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. p. 138, Figure 8.15.B, right panel PhysicianBench (pass@1); orange Sonnet 5.5 low label (visually read printed labels).
    Sonnet 5.5 | PhysicianBench (pass@1) | low 27.2%
    Every result from this document

Other Anthropic models: Claude 3.7 Sonnet, Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.1, Claude Opus 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 4, Claude Sonnet 4.5, Claude Sonnet 4.6, Claude Sonnet 5