Clinical Benchmarks

Claude Opus 5.5

Released 22 Sep 2026$4 input, $20 output per million tokens1M contextproprietaryAlso written as anthropic/claude-opus-5-5, claude-opus-5-5Compare with other models

Clinical Benchmarks Index
88.8rank 2 of 148; 88.8 × 1 = 88.8, from 4 of 10 boards
Boards
4 of 13
Results
4
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. HealthBench Professional
    Anthropic; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt; safety classifiers and refusal fallback to Opus 5; length-adjusted; raw 77.1%
    0.656
    Rank 4 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Sep 2026

Documentation and coding

  1. MedScribe (Vals AI)
    model ID anthropic/claude-opus-5-5; temperature=1; max_output_tokens=128000; compute_effort=max
    91.43
    Rank 1 of 105Leads this board
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000
    49.80
    Rank 16 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    68.4
    Rank 1 of 21 here, 12 models on the boardLeads this board
    Vendor-reported
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 4HealthBench Professional Anthropic; adaptive max effort; Opus 4.8 grader; five trials; no tools or custo… 0.656
    Printed as 65.6%Vendor-reported, measured Sep 2026Configuration: Anthropic; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt; safety classifiers and refusal fallback to Opus 5; length-adjusted; raw 77.1%
    Claude Opus 5.5 System Card system card, Anthropic, 22 Sep 2026. pp. 213–214, Section 8.15.2, paragraph and Figure 8.15.2.A.
    8.15.2 HealthBench Professional results length adjustment, which penalizes verbose model responses, Claude Opus 5.5 achieved a score of 65.6%.
    Every result from this document
  2. 1MedScribe (Vals AI) model ID anthropic/claude-opus-5-5; temperature=1; max_output_tokens=128000; co… 91.43
    Printed as 91.43%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-5-5; temperature=1; max_output_tokens=128000; compute_effort=max
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 1 of 105 (Claude Opus 5.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-5-5"].
    1 | Claude Opus 5.5 | 91.43%±1.93 | $4/$20 | 7m22s
    Every result from this document
  3. 16MedCode (Vals AI) model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_outp… 49.80
    Printed as 49.80%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 16 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-5-5"].
    16 | Claude Opus 5.5 | 49.80%±2.27 | $4/$20 | 4m07s
    Every result from this document
  4. 1PhysicianBench Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 68.4
    Printed as 68.4%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Opus 5.5 (max).
    PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
    Every result from this document

Other Anthropic models: Claude 3.7 Sonnet, Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.1, Claude Opus 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Claude Sonnet 4, Claude Sonnet 4.5, Claude Sonnet 4.6, Claude Sonnet 5, Claude Sonnet 5.5