DeepSeek R1
Released 20 Jan 2025128K context671B total / 37B activemitCompare with other models
- Clinical Benchmarks Index
- 18.8rank 129 of 148; 23.1 × 0.816 = 18.8, from 2 of 10 boards
- Boards
- 2 of 13
- Results
- 2
- Latest measurement
- Aug 2026
Results
Each line is placed on its own board. Dark tick: the board leader.
Clinical reasoning and knowledge
- 0.485Rank 7 of 10 here, 11 models on the boardLeader Gemini 3.1 Pro (Preview) 0.652
Safety
- 55.8Rank 17 of 17 here, 19 models on the boardLeader LiSA 2.5 86.2
Sources
Open a line for the quote and page.
7MedHELM 0.485
Printed as 0.485Official leaderboard, measured May 2026MedHELM leaderboard (medhelm.org), v5.0.0 official leaderboard, Stanford CRFM (MedHELM), 14 May 2026. medhelm.org home, 'Current leaders Mean win rate v5.0.0' table ('10 of 11 models · Updated 14 May 2026')7 DeepSeek R1 DeepSeek 0.485
Every result from this document- Also printed in MedHELM v5.0.0 release data: medhelm_scenarios group table (mean win rate): 0.48541666666666666, releases/v5.0.0/groups/medhelm_scenarios.json, table 'Accuracy', column 'Mean win rate'
17First, Do NOHARM (v2) 55.8
Printed as 55.8%Official leaderboard, measured Aug 2026MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)16DeepSeek R1OSSDeepSeek 55.8%
Every result from this document
Other DeepSeek models: DeepSeek-V3.1, DeepSeek-V3.2-Exp, DeepSeek V4.1 Flash, DeepSeek V4 Flash, DeepSeek V4 Flash 0731, DeepSeek V4 Pro