Clinical Benchmarks

Mercury 2.5

Released 8 Sep 2026$0.2 input, $0.75 output per million tokens260K contextproprietaryAlso written as inception/mercury-2.5Compare with other models

Clinical Benchmarks Index
34.6rank 106 of 148; 42.4 × 0.816 = 34.6, from 2 of 10 boards
Boards
2 of 13
Results
2
Latest measurement
Sep 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Documentation and coding

  1. MedScribe (Vals AI)
    model ID inception/mercury-2.5; temperature=1; max_output_tokens=65536; reasoning_effort=high
    55.09
    Rank 102 of 105Leader Claude Opus 5.5 91.43
    Official leaderboard
    Measured Sep 2026
  2. MedCode (Vals AI)
    model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_output_tokens=65536
    31.33
    Rank 91 of 103Leader Claude Opus 5 63.57
    Official leaderboard
    Measured Sep 2026

Sources

Open a line for the quote and page.

  1. 102MedScribe (Vals AI) model ID inception/mercury-2.5; temperature=1; max_output_tokens=65536; reasoni… 55.09
    Printed as 55.09%Official leaderboard, measured Sep 2026Configuration: model ID inception/mercury-2.5; temperature=1; max_output_tokens=65536; reasoning_effort=high
    Vals AI MedScribe leaderboard official leaderboard, Vals AI, 3 Sep 2026. Vals AI MedScribe leaderboard, View: All Models, Task: Overall, row 102 of 105 (Mercury 2.5), Accuracy column; Updated 9/29/2026. Rendered BenchmarkView table; configuration from embedded astro-island BenchmarkView props, benchmarkView.default.tasks.overall["inception/mercury-2.5"].
    102 | Mercury 2.5 | 55.09%±2.10 | $0.2/$0.75 | 9.99s
    Every result from this document
  2. 91MedCode (Vals AI) model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_outpu… 31.33
    Printed as 31.33%Official leaderboard, measured Sep 2026Configuration: model ID inception/mercury-2.5; reasoning_effort=high; temperature=1; max_output_tokens=65536
    Vals AI MedCode leaderboard official leaderboard, Vals AI, 29 Sep 2026. MedCode leaderboard, Overall task, All Models expanded, rank 91 of 103; Accuracy column; board Updated 9/29/2026. Rendered table row; configuration from embedded BenchmarkView props benchmarkView.default.tasks.overall["inception/mercury-2.5"].
    91 | Mercury 2.5 | 31.33%±1.95 | $0.2/$0.75 | 7.93s
    Every result from this document