Clinical Benchmarks

Anthropic

Anthropic models with results here, each placed on its own board.

Models
16
Results
92
Boards
13 of 13

www.anthropic.com

Models

ModelKindReleasedBoards
Claude Sonnet 5.5Model28 Sep 20264
Claude Opus 5.5Model22 Sep 20264
Claude Fable 5.1Model1 Sep 20264
Claude Opus 5Model24 Jul 20268
Claude Sonnet 5Model30 Jun 20266
Claude Fable 5Model9 Jun 20266
Claude Opus 4.8Model28 May 20266
Claude Opus 4.7Model16 Apr 20266
Claude Sonnet 4.6Model17 Feb 20266
Claude Opus 4.6Model5 Feb 20268
Claude Opus 4.5Model24 Nov 20253
Claude Haiku 4.5Model15 Oct 20253
Claude Sonnet 4.5Model29 Sep 20252
Claude Opus 4.1Model5 Aug 20252
Claude Sonnet 4Model22 May 20252
Claude 3.7 SonnetModel24 Feb 20251

Results

One line per result. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. HealthBench Professional
    Anthropic; max effort; Opus 4.8 grader; safety classifiers enabled; length-adjusted; raw 77.1%
    0.692
    Rank 2 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Sep 2026
  2. HealthBench Professional
    length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 70.3%). Measured as Claude Mythos 5; Anthropic itself prints 66.0 in the Fable 5 column of the Opus 5 card with footnote Mythos 5.
    0.660
    Rank 3 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026
  3. HealthBench Professional
    Anthropic; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt; safety classifiers and refusal fallback to Opus 5; length-adjusted; raw 77.1%
    0.656
    Rank 4 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Sep 2026
  4. HealthBench Professional
    Measured as Claude Fable 5; Anthropic September card; length-adjusted; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt; raw 68.9%
    0.633
    Rank 6 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Sep 2026
  5. HealthBench Professional
    length-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%).
    0.621
    Rank 7 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Sep 2026
  6. HealthBench Professional
    length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%).
    0.598
    Rank 11 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jul 2026
  7. HealthBench Professional
    length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).
    0.578
    Rank 13 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026
  8. HealthBench Professional
    Anthropic June evaluation; length-adjusted; adaptive max effort; Opus 4.8 grader; five trials; no tools or custom system prompt
    0.574
    Rank 15 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026
  9. HealthBench Professional
    length-adjusted, adaptive thinking at max effort, Claude Sonnet 4.6 grader (the grader used by the Opus 4.8 card itself); the other Claude rows on this board use the Claude Opus 4.8 grader, under which Anthropic later prints 57.4 for this model.
    0.558
    Rank 17 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured May 2026
  10. HealthBench Professional
    length-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 trials; comparison model in the Opus 4.8 card
    0.519
    Rank 21 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured May 2026
  11. HealthBench Professional
    length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Sonnet 5 card. Other Anthropic prints: 44.4% (Fable card Figure 8.18.2.A), 41.7% (Opus 4.8 card, Sonnet 4.6 grader)
    0.442
    Rank 27 of 31Leader GPT-6 Astra (Anthropic run) 0.703
    Vendor-reported
    Measured Jun 2026
  12. MedXpertQA (MM)
    Qwen-run comparison in the Qwen3.8-Max launch post
    80.0
    Rank 4 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Independent run
  13. MedXpertQA (MM)
    Qwen-run comparison in the Qwen3.8-Max launch post
    71.7
    Rank 9 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Independent run
  14. MedXpertQA (MM)
    Meta's Muse Spark launch table (Meta reports the better of vendor self-reports and its own reproduction)
    64.8
    Rank 15 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Independent run
    Measured Apr 2026
  15. MedXpertQA (MM)
    Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
    63.6
    Rank 16 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Independent run
    Measured Feb 2026
  16. 57.1
    Rank 6 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
    Official leaderboard
    Measured Aug 2026
  17. 56.6
    Rank 7 of 8 here, 11 models on the boardLeader GPT-5.6 Sol 60.2
    Official leaderboard
    Measured Aug 2026
  18. 0.456
    Rank 8 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
    Official leaderboard
    Measured May 2026
  19. 0.450
    Rank 9 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
    Official leaderboard
    Measured May 2026

Documentation and coding

  1. MedScribe (Vals AI)
    model ID anthropic/claude-opus-5-5; temperature=1; max_output_tokens=128000; compute_effort=max
    91.43
    Rank 1 of 105Leads this board
    Official leaderboard
    Measured Sep 2026
  2. 91.29
    Rank 2 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  3. MedScribe (Vals AI)
    model ID anthropic/claude-sonnet-5-5; temperature=1; max_output_tokens=128000; compute_effort=max
    91.10
    Rank 3 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  4. 90.98
    Rank 4 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  5. 88.52
    Rank 10 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  6. 86.74
    Rank 18 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  7. MedScribe (Vals AI)
    model ID anthropic/claude-opus-4-6-thinking; temperature=1; max_output_tokens=30000; compute_effort=max
    86.13
    Rank 21 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  8. MedScribe (Vals AI)
    model ID anthropic/claude-opus-4-8; temperature=1; max_output_tokens=30000; compute_effort=max
    85.75
    Rank 23 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  9. MedScribe (Vals AI)
    model ID anthropic/claude-opus-4-5-20251101-thinking; temperature=1; max_output_tokens=30000; compute_effort=high
    85.32
    Rank 26 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  10. MedScribe (Vals AI)
    model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000
    85.23
    Rank 28 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  11. MedScribe (Vals AI)
    model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000
    84.52
    Rank 31 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  12. MedScribe (Vals AI)
    model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000
    84.10
    Rank 36 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  13. MedScribe (Vals AI)
    model ID anthropic/claude-opus-4-5-20251101; temperature=1; max_output_tokens=30000; compute_effort=high
    83.25
    Rank 44 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  14. MedScribe (Vals AI)
    model ID anthropic/claude-opus-4-7; temperature=1; max_output_tokens=30000; compute_effort=max
    82.95
    Rank 46 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  15. MedScribe (Vals AI)
    model ID anthropic/claude-sonnet-5; temperature=1; max_output_tokens=30000; compute_effort=max
    76.05
    Rank 74 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  16. MedScribe (Vals AI)
    model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000
    73.90
    Rank 79 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  17. MedScribe (Vals AI)
    model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000
    72.41
    Rank 84 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  18. MedScribe (Vals AI)
    model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000
    71.75
    Rank 88 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  19. MedScribe (Vals AI)
    model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000
    69.35
    Rank 92 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  20. 63.57
    Rank 1 of 103Leads this board
    Official leaderboard
    Measured Sep 2026
  21. 56.07
    Rank 3 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  22. 54.86
    Rank 6 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  23. 53.51
    Rank 7 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  24. 53.22
    Rank 9 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  25. MedCode (Vals AI)
    model ID anthropic/claude-sonnet-5-5; compute_effort=max; temperature=1; max_output_tokens=128000
    52.92
    Rank 11 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  26. MedCode (Vals AI)
    model ID anthropic/claude-opus-5-5; compute_effort=max; temperature=1; max_output_tokens=128000
    49.80
    Rank 16 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  27. MedCode (Vals AI)
    model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000
    49.16
    Rank 21 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  28. MedCode (Vals AI)
    model ID anthropic/claude-opus-4-6-thinking; compute_effort=max; temperature=1; max_output_tokens=30000
    49.13
    Rank 22 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  29. MedCode (Vals AI)
    model ID anthropic/claude-opus-4-6; compute_effort=max; temperature=1; max_output_tokens=30000
    48.24
    Rank 27 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  30. MedCode (Vals AI)
    model ID anthropic/claude-sonnet-5; compute_effort=max; temperature=1; max_output_tokens=30000
    47.54
    Rank 30 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  31. MedCode (Vals AI)
    model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000
    47.23
    Rank 32 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  32. MedCode (Vals AI)
    model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000
    45.17
    Rank 35 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  33. MedCode (Vals AI)
    model ID anthropic/claude-sonnet-4-5-20250929-thinking; temperature=1; max_output_tokens=30000
    44.13
    Rank 39 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  34. MedCode (Vals AI)
    model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000
    41.37
    Rank 51 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  35. MedCode (Vals AI)
    model ID anthropic/claude-sonnet-4-5-20250929; temperature=1; max_output_tokens=30000
    40.57
    Rank 59 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  36. MedCode (Vals AI)
    model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000
    34.96
    Rank 75 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  37. MedCode (Vals AI)
    model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000
    33.94
    Rank 79 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  38. MedCode (Vals AI)
    model ID anthropic/claude-haiku-4-5-20251001-thinking; temperature=1; max_output_tokens=30000
    32.68
    Rank 84 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

EHR and workflow agents

  1. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    68.4
    Rank 1 of 21 here, 12 models on the boardLeads this board
    Vendor-reported
    Measured Sep 2026
  2. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    63.2
    Rank 2 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  3. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    61.0
    Rank 3 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  4. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    57.6
    Rank 4 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  5. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; xhigh effort; Opus 5 rubric grader; safety classifiers enabled
    56.4
    Rank 5 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  6. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; high effort; Opus 5 rubric grader; safety classifiers enabled
    47.6
    Rank 6 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  7. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; max effort; Opus 5 rubric grader; safety classifiers enabled
    37.4
    Rank 8 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  8. 31.7
    Rank 9 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Official leaderboard
    Measured May 2026
  9. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; medium effort; Opus 5 rubric grader; safety classifiers enabled
    30.0
    Rank 10 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  10. 29.3
    Rank 11 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Official leaderboard
    Measured May 2026
  11. PhysicianBench
    Anthropic harness; 100 tasks; pass@1; low effort; Opus 5 rubric grader; safety classifiers enabled
    27.2
    Rank 13 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Vendor-reported
    Measured Sep 2026
  12. 23.0
    Rank 14 of 21 here, 12 models on the boardLeader Claude Opus 5.5 (max) 68.4
    Official leaderboard
    Measured May 2026
  13. EHR-Complex
    validation configuration
    0.360
    Rank 13 of 18Leader GPT-5.4 (high reasoning) 0.650
    Official leaderboard
    Measured Jun 2026
  14. HealthAgentBench
    $3.3/task; harness+model evaluated jointly
    55
    Rank 1 of 12Leads this board
    Official leaderboard
    Measured Jul 2026
  15. 36
    Rank 4 of 12Leader Claude Code (Opus 5) 55
    Official leaderboard
    Measured Jul 2026
  16. 32
    Rank 6 of 12Leader Claude Code (Opus 5) 55
    Official leaderboard
    Measured Jul 2026
  17. 27
    Rank 8 of 12Leader Claude Code (Opus 5) 55
    Official leaderboard
    Measured Jul 2026
  18. 19
    Rank 10 of 12Leader Claude Code (Opus 5) 55
    Official leaderboard
    Measured Jul 2026
  19. 17
    Rank 11 of 12Leader Claude Code (Opus 5) 55
    Official leaderboard
    Measured Jul 2026
  20. CHI-Bench
    community-submitted harness config validated by automated workspace judge
    54.7
    Rank 1 of 43 here, 45 models on the boardLeads this board
    Independent run
    Measured Aug 2026
  21. 37.3
    Rank 2 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Jul 2026
  22. 37.3
    Rank 2 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Independent run
    Measured Aug 2026
  23. 33.3
    Rank 4 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  24. 28.0
    Rank 5 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  25. 26.2
    Rank 6 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Apr 2026
  26. 24.4
    Rank 9 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Apr 2026
  27. 24.0
    Rank 10 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Jul 2026
  28. 20.0
    Rank 13 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Jul 2026
  29. CHI-Bench
    All Domains pass@1; openclaw harness
    17.3
    Rank 17 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured May 2026
  30. CHI-Bench
    All Domains pass@1; claude-code harness
    6.2
    Rank 36 of 43 here, 45 models on the boardLeader erius + claude-opus-5 54.7
    Official leaderboard
    Measured Apr 2026
  31. HealthAdminBench
    screenshot-only, detailed prompting; native CUA harness; subtask rate 78.4%
    36.3
    Rank 1 of 7 here, 5 models on the boardLeads this board
    Official leaderboard
    Measured Apr 2026
  32. HealthAdminBench
    screenshot-only, detailed prompting; authors' standardized harness, no native CUA
    14.8
    Rank 4 of 7 here, 5 models on the boardLeader Claude Opus 4.6 (computer-use agent) 36.3
    Official leaderboard
    Measured Apr 2026

Safety

  1. MedPIC
    zero-shot; independent questions; exact option-set match
    64.2
    Rank 9 of 28Leader Gemini-3.1-Pro 80.7
    Independent run
    Measured Aug 2026
  2. First, Do NOHARM (v2)
    v2 run on ARISE; 19 models on the board
    74.6
    Rank 6 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026
  3. 65.0
    Rank 11 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026