Clinical Benchmarks

Claude Opus 4.5

Released 24 Nov 2025$5 input, $25 output per million tokensproprietaryAlso written as anthropic/claude-opus-4-5-20251101, Claude Opus 4.5 (Thinking), Claude Opus 4.5 (Nonthinking), anthropic/claude-opus-4-5-20251101-thinking, claude-opus-4-5-20251101, claude-opus-4-5-20251101-thinkingCompare with other models

Clinical Benchmarks Index
76.4rank 15 of 148; 76.4 × 1 = 76.4, from 3 of 10 boards
Boards
3 of 13
Results
5
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. MedXpertQA (MM)
    Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
    63.6
    Rank 16 of 22 here, 5 models on the boardLeader GPT-5.6 Sol 81.5
    Independent run
    Measured Feb 2026

Documentation and coding

  1. MedScribe (Vals AI)
    model ID anthropic/claude-opus-4-5-20251101-thinking; temperature=1; max_output_tokens=30000; compute_effort=high
    85.32
    Rank 26 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedScribe (Vals AI)
    model ID anthropic/claude-opus-4-5-20251101; temperature=1; max_output_tokens=30000; compute_effort=high
    83.25
    Rank 44 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  3. MedCode (Vals AI)
    model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000
    49.16
    Rank 21 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026
  4. MedCode (Vals AI)
    model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000
    45.17
    Rank 35 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 26MedScribe (Vals AI) model ID anthropic/claude-opus-4-5-20251101-thinking; temperature=1; max_output… 85.32
    Printed as 85.32%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-5-20251101-thinking; temperature=1; max_output_tokens=30000; compute_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 26 of 105 (Claude Opus 4.5 (Thinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-5-20251101-thinking"].
    26 | Claude Opus 4.5 (Thinking) | 85.32%±1.90 | $5/$25 | 72.11s
    Every result from this document
  2. 44MedScribe (Vals AI) model ID anthropic/claude-opus-4-5-20251101; temperature=1; max_output_tokens=3… 83.25
    Printed as 83.25%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-5-20251101; temperature=1; max_output_tokens=30000; compute_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 44 of 105 (Claude Opus 4.5 (Nonthinking)), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["anthropic/claude-opus-4-5-20251101"].
    44 | Claude Opus 4.5 (Nonthinking) | 83.25%±1.93 | $5/$25 | 43.30s
    Every result from this document
  3. 21MedCode (Vals AI) model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temp… 49.16
    Printed as 49.16%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-5-20251101-thinking; compute_effort=high; temperature=1; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 21 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-4-5-20251101-thinking"].
    21 | Claude Opus 4.5 (Thinking) | 49.16%±2.01 | $5/$25 | 60.84s
    Every result from this document
  4. 35MedCode (Vals AI) model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1… 45.17
    Printed as 45.17%Official leaderboard, measured Sep 2026Configuration: model ID anthropic/claude-opus-4-5-20251101; compute_effort=high; temperature=1; max_output_tokens=30000
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 35 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["anthropic/claude-opus-4-5-20251101"].
    35 | Claude Opus 4.5 (Nonthinking) | 45.17%±1.89 | $5/$25 | 5.02s
    Every result from this document
  5. 16MedXpertQA (MM) Qwen3.5 model card comparison; source-specific evaluation, not harmonized acros… 63.6
    Printed as 63.6Independent run, measured Feb 2026Configuration: Qwen3.5 model card comparison; source-specific evaluation, not harmonized across vendors
    Qwen/Qwen3.5-397B-A17B model card model card, Alibaba / Qwen, 16 Feb 2026. Benchmark Results > Vision Language > Medical VQA; MedXpertQA-MM row, Claude 4.5 Opus column.
    | GPT5.2 | Claude 4.5 Opus | Gemini-3 Pro | Qwen3-VL-235B-A22B | K2.5-1T-A32B | Qwen3.5-397B-A17B MedXpertQA-MM | 73.3 | 63.6 | 76.0 | 47.6 | 65.3 | 70.0
    Every result from this document

Other Anthropic models: Claude 3.7 Sonnet, Claude Fable 5, Claude Fable 5.1, Claude Haiku 4.5, Claude Opus 4.1, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Claude Opus 5.5, Claude Sonnet 4, Claude Sonnet 4.5, Claude Sonnet 4.6, Claude Sonnet 5, Claude Sonnet 5.5