Clinical Benchmarks

Claude Opus 5

Released 24 Jul 2026$5 input, $25 output per million tokens1.0M contextproprietaryAlso written as Claude Opus 5 (Adaptive Reasoning, Max Effort), claude-opus-5, Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Compare with other models

Clinical Benchmarks Index
80.6rank 11 of 148; 80.6 × 1 = 80.6, from 6 of 10 boards
Boards
8 of 13
Results
9
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. HealthBench Professional
    length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%).
    0.598
    Rank 11 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jul 2026
  2. 57.1
    Rank 6 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
    Official leaderboard
    Measured Aug 2026

Documentation and coding

  1. 90.98
    Rank 4 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. 63.57
    Rank 1 of 103Leads this board
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    57.6
    Rank 4 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  2. HealthAgentBench
    Listed as Claude Code (Opus 5)
    $3.3/task; harness+model evaluated jointly
    55
    Rank 1 of 12Leads this board
    Official leaderboard
    Measured Jul 2026
  3. CHI-Bench
    Listed as erius + claude-opus-5
    community-submitted harness config validated by automated workspace judge
    54.7
    Rank 1 of 43 here, 45 models on the boardLeads this board
    Independent run
    Measured Aug 2026
  4. CHI-Bench
    Listed as claude-code + claude-opus-5
    37.3
    Rank 2 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Jul 2026

Safety

  1. First, Do NOHARM (v2)
    v2 run on ARISE; 19 models on the board
    74.6
    Rank 6 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026

Sources

Open a line for the quote and page.

  1. 11HealthBench Professional length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Op… 0.598
    Printed as 59.8%Vendor-reported, measured Jul 2026Configuration: length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%).
    System Card: Claude Opus 5 system card, Anthropic, 24 Jul 2026. p. 189, section 8.15.2 HealthBench Professional results; also Table 8.1.A p. 152 ('HealthBench Professional 59.8 ...')
    Claude Opus 5 achieved a raw score of 73.4%, which is the highest amongst all Claude models, ahead of Claude Mythos 5 at 70.3%, Claude Opus 4.8 at 60.3%, and Claude Sonnet 5 at 62.4%. After length adjustment, which penalizes verbose model responses, Claude Opus 5 achieved a score of 59.8%.
    Every result from this document
  2. 4MedScribe (Vals AI) 90.98
    Printed as 90.98%Official leaderboard, measured Sep 2026
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 4 of 105 (Claude Opus 5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-5"].
    4 | Claude Opus 5 | 90.98%±1.92 | $5/$25 | 76.56s
    Every result from this document
  3. 1MedCode (Vals AI) 63.57
    Printed as 63.57%Official leaderboard, measured Sep 2026
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 1 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-5"].
    1 | Claude Opus 5 | 63.57%±1.99 | $5/$25 | 24.50s
    Every result from this document
  4. 6MAST (Medical AI Superintelligence Test) 57.1
    Printed as 57.1%Official leaderboard, measured Aug 2026
    MAST: Medical AI Superintelligence Test leaderboard (General board) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast, 'Which AI can you trust for medical questions?' General tab, composite score table (8 of 11 models shown; 'Last updated August 15, 2026')
    6 Claude Opus 5 Anthropic 57.1%
    Every result from this document
  5. 6First, Do NOHARM (v2) v2 run on ARISE; 19 models on the board 74.6
    Printed as 74.6%Official leaderboard, measured Aug 2026Configuration: v2 run on ARISE; 19 models on the board
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    6Claude Opus 5Anthropic 74.6%
    Every result from this document
  6. 4PhysicianBench Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety… 57.6
    Printed as 57.6%Vendor-reported, measured Sep 2026Configuration: Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    Claude Sonnet 5.5 System Card system card, Anthropic, 28 Sep 2026. pp. 138–139, Section 8.15.3; PhysicianBench pass@1, Claude Opus 5 (max).
    PhysicianBench is a public benchmark of 100 physician tasks carried out in an EHR. At max effort, Claude Sonnet 5.5 passes 63.2% of attempts, well above Claude Sonnet 5 (37.4%), level with Claude Fable 5.1 (61.0%), about 6 points above Claude Opus 5 (57.6%) and about 5 points below Claude Opus 5.5 (68.4%).
    Every result from this document
  7. 1HealthAgentBench $3.3/task; harness+model evaluated jointly 55
    Printed as 55%Official leaderboard, measured Jul 2026Configuration: $3.3/task; harness+model evaluated jointly
    HealthAgentBench leaderboard official leaderboard, Microsoft Research (HealthAgentBench), 27 Jul 2026. Leaderboard table (homepage), rank 1, Success Rate column
    1 Claude Code (Opus 5) 55% $3.3
    Every result from this document
  8. 1CHI-Bench community-submitted harness config validated by automated workspace judge 54.7
    Printed as 54.7%Independent run, measured Aug 2026Configuration: community-submitted harness config validated by automated workspace judge
    CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 01, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-07-26; accessed 2026-09-30.
    01 | erius submitted by Michael Johnson (MJ) | claude-opus-5 | Proprietary | 54.7% | 72.0% | 36.0% | 56.0% | 2026-07-26
    Every result from this document
  9. 2CHI-Bench 37.3
    Printed as 37.3%Official leaderboard, measured Jul 2026
    CHI-Bench leaderboard (actAVA) official leaderboard, actAVA, 12 Aug 2026. CHI-Bench v1.0.0, All Domains, rank 03, Agent/Model and Accuracy columns; PA/UM/CM follow; board last updated 2026-08-12; board Date column 2026-07-24; accessed 2026-09-30. Linked submission provenance: run 2026-07-24T19:08:02Z to 2026-07-24T23:22:36Z; submitted_at 2026-08-12T20:20:00Z.
    03 | claude-code | claude-opus-5 | Proprietary | 37.3% | 20.0% | 32.0% | 60.0% | 2026-07-24
    Every result from this document

Other Anthropic models: Claude 3.7 Sonnet, Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.1, Claude Opus 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5.5, Claude Sonnet 4, Claude Sonnet 4.5, Claude Sonnet 4.6, Claude Sonnet 5, Claude Sonnet 5.5