Clinical Benchmarks

Claude Sonnet 4

Released 22 May 2025proprietaryAlso written as claude-sonnet-4-20250514-thinking, anthropic/claude-sonnet-4-20250514-thinking, claude-sonnet-4-20250514, Claude Sonnet 4 (Thinking), Claude Sonnet 4 (Nonthinking), anthropic/claude-sonnet-4-20250514Compare with other models

Clinical Benchmarks Index
46.1rank 73 of 148; 56.5 × 0.816 = 46.1, from 2 of 10 boards
Boards
2 of 13
Results
4
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Documentation and coding

  1. MedScribe (Vals AI)
    model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000
    72.41
    Rank 84 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedScribe (Vals AI)
    model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000
    69.35
    Rank 92 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  3. MedCode (Vals AI)
    model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000
    34.96
    Rank 75 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  4. MedCode (Vals AI)
    model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000
    33.94
    Rank 79 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 84MedScribe (Vals AI) model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=3… 72.41
    Printed as 72.41%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 84 of 105 (Claude Sonnet 4 (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-20250514"].
    84 | Claude Sonnet 4 (Nonthinking) | 72.41%±1.93 | $3/$15 | 25.67s
    Every result from this document
  2. 92MedScribe (Vals AI) model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000 69.35
    Printed as 69.35%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 92 of 105 (Claude Sonnet 4 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-20250514-thinking"].
    92 | Claude Sonnet 4 (Thinking) | 69.35%±2.21 | $3/$15 | 39.57s
    Every result from this document
  3. 75MedCode (Vals AI) model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000 34.96
    Printed as 34.96%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-20250514-thinking; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 75 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-20250514-thinking"].
    75 | Claude Sonnet 4 (Thinking) | 34.96%±1.94 | $3/$15 | 39.80s
    Every result from this document
  4. 79MedCode (Vals AI) model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=3… 33.94
    Printed as 33.94%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-sonnet-4-20250514; temperature=1; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 79 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-sonnet-4-20250514"].
    79 | Claude Sonnet 4 (Nonthinking) | 33.94%±1.91 | $3/$15 | 7.30s
    Every result from this document

Other Anthropic models: Claude 3.7 Sonnet, Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.1, Claude Opus 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 4.5, Claude Sonnet 4.6, Claude Sonnet 5, Claude Sonnet 5.5