Clinical Benchmarks

Claude Opus 4.1

Released 5 Aug 2025proprietaryAlso written as anthropic/claude-opus-4-1-20250805-thinking, claude-opus-4-1-20250805, anthropic/claude-opus-4-1-20250805, claude-opus-4-1-20250805-thinking, Claude Opus 4.1 (Thinking), Claude Opus 4.1 (Nonthinking)Compare with other models

Clinical Benchmarks Index
58.2rank 47 of 148; 71.3 × 0.816 = 58.2, from 2 of 10 boards
Boards
2 of 13
Results
4
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Documentation and coding

  1. MedScribe (Vals AI)
    model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000
    73.90
    Rank 79 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedScribe (Vals AI)
    model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000
    71.75
    Rank 88 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  3. MedCode (Vals AI)
    model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000
    47.23
    Rank 32 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  4. MedCode (Vals AI)
    model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000
    41.37
    Rank 51 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 79MedScribe (Vals AI) model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output… 73.90
    Printed as 73.90%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 79 of 105 (Claude Opus 4.1 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-1-20250805-thinking"].
    79 | Claude Opus 4.1 (Thinking) | 73.90%±1.97 | $15/$75 | 57.40s
    Every result from this document
  2. 88MedScribe (Vals AI) model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=3… 71.75
    Printed as 71.75%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 88 of 105 (Claude Opus 4.1 (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-1-20250805"].
    88 | Claude Opus 4.1 (Nonthinking) | 71.75%±2.02 | $15/$75 | 38.04s
    Every result from this document
  3. 32MedCode (Vals AI) model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output… 47.23
    Printed as 47.23%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-1-20250805-thinking; temperature=1; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 32 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-4-1-20250805-thinking"].
    32 | Claude Opus 4.1 (Thinking) | 47.23%±2.07 | $15/$75 | 33.26s
    Every result from this document
  4. 51MedCode (Vals AI) model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=3… 41.37
    Printed as 41.37%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-1-20250805; temperature=1; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 51 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-4-1-20250805"].
    51 | Claude Opus 4.1 (Nonthinking) | 41.37%±1.96 | $15/$75 | 13.08s
    Every result from this document

Other Anthropic models: Claude 3.7 Sonnet, Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 4, Claude Sonnet 4.5, Claude Sonnet 4.6, Claude Sonnet 5, Claude Sonnet 5.5