Clinical Benchmarks

DeepSeek R1

Released 20 Jan 2025128K context671B total / 37B activemitCompare with other models

Clinical Benchmarks Index
18.8rank 129 of 148; 23.1 × 0.816 = 18.8, from 2 of 10 boards
Boards
2 of 13
Results
2
Latest measurement
Aug 2026

Results

Each line is placed on its own board. Dark tick: the board leader.

Clinical reasoning and knowledge

  1. 0.485
    Rank 7 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
    Official leaderboard
    Measured May 2026

Safety

  1. 55.8
    Rank 17 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
    Official leaderboard
    Measured Aug 2026

Sources

Open a line for the quote and page.

  1. 7MedHELM 0.485
    Printed as 0.485Official leaderboard, measured May 2026
    MedHELM leaderboard (medhelm.org), v5.0.0 official leaderboard, Stanford CRFM (MedHELM), 14 May 2026. medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')
    7 DeepSeek R1 DeepSeek 0.485
    Every result from this document
  2. 17First, Do NOHARM (v2) 55.8
    Printed as 55.8%Official leaderboard, measured Aug 2026
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    16DeepSeek R1OSSDeepSeek 55.8%
    Every result from this document

Other DeepSeek models: DeepSeek-V3.1, DeepSeek-V3.2-Exp, DeepSeek V4.1 Flash, DeepSeek V4 Flash, DeepSeek V4 Flash 0731, DeepSeek V4 Pro