First, Do NOHARM (v2)
How often, and how severely, model consultation recommendations contain potentially harmful errors.
More about this boardLess
Frequency and severity of potentially harmful errors in LLM-generated medical consultation recommendations (Numerous Options Harm Assessment for Risk in Medicine); primary-care-to-specialist consults.
Paper: arXiv 2512.01241. The v1 study found potential for severe harm in up to 24.6 percent of directly applied recommendations, with errors of omission behind more than 80 percent of the severe cases.
Published by Stanford/Harvard-led consortium (50+ researchers incl. 29 board-certified physicians); hosted by ARISE, released Dec 2025. 1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options. Percentage safety score, higher better.
- Headline metric
- F1 (Weighted)
- Board marked as preview
- Yes
- Board last updated
- 15 Aug 2026
- Models in the full dataset
- 43
- Models in the board's default view
- 19
- Rows
- 17
- Models
- 17 of 19on the board
- Labs
- 12
- Last measured
- Sep 2026
- Leader
- 86.2LiSA 2.5
Ranking
Percentage safety score, higher better
- Anthropic
- OpenAI
- Alibaba
- Meta
- Moonshot AI
- Other labs
17 of 17 rows
Rows and sources
Open a row for the quote, the page and the document.
1LiSA 2.5 First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublished 86.2
Printed as 86.2%Official leaderboard, measured Sep 2026Configuration: First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublishedARISE MAST technical leaderboard official leaderboard, ARISE. First Do NOHARM v2, overall leaderboard; LiSA 2.5 row; displayed 2026-09-30.First Do NOHARM v2 overall metric across 19 models 1LiSA 2.5RAGAMBOSS 86.2%
Every result from this document2Doximity Ask 6.1 First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublished 84.5
Printed as 84.5%Official leaderboard, measured Sep 2026Configuration: First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublishedARISE MAST technical leaderboard official leaderboard, ARISE. First Do NOHARM v2, overall leaderboard; Doximity Ask 6.1 row; displayed 2026-09-30.First Do NOHARM v2 overall metric across 19 models 2Doximity Ask 6.1RAGDoximity 84.5%
Every result from this document3OpenEvidence First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublished 80.0
Printed as 80.0%Official leaderboard, measured Sep 2026Configuration: First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublishedARISE MAST technical leaderboard official leaderboard, ARISE. First Do NOHARM v2, overall leaderboard; OpenEvidence row; displayed 2026-09-30.First Do NOHARM v2 overall metric across 19 models 3OpenEvidenceRAGOpenEvidence 80.0%
Every result from this document4Glass 5.6 Max First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublished 79.7
Printed as 79.7%Official leaderboard, measured Sep 2026Configuration: First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublishedARISE MAST technical leaderboard official leaderboard, ARISE. First Do NOHARM v2, overall leaderboard; Glass 5.6 Max row; displayed 2026-09-30.First Do NOHARM v2 overall metric across 19 models 4Glass 5.6 MaxRAGGlass Health 79.7%
Every result from this document4Muse Spark 1.1 79.7
Printed as 79.7%Official leaderboard, measured Aug 2026MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)5Muse Spark 1.1Meta 79.7%
Every result from this document6Claude Opus 5 v2 run on ARISE; 19 models on the board 74.6
Printed as 74.6%Official leaderboard, measured Aug 2026Configuration: v2 run on ARISE; 19 models on the boardMAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column6Claude Opus 5Anthropic 74.6%
Every result from this document7Kimi K3 74.0
Printed as 74.0%Official leaderboard, measured Aug 2026MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column7Kimi K3OSSMoonshot AI 74.0%
Every result from this document8GPT-5.6 Sol 70.1
Printed as 70.1%Official leaderboard, measured Aug 2026MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column8GPT-5.6 SolOpenAI 70.1%
Every result from this document9GPT-5.5 70.0
Printed as 70.0%Official leaderboard, measured Aug 2026MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column9GPT-5.5OpenAI 70.0%
Every result from this document10GPT-5 from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI)… 68.6
Printed as 68.6%Official leaderboard, measured Aug 2026Configuration: from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI); not in the Latest Flagships rankingMAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, Model Leaderboard (Top 10 shown), row 10, SAFETY column10 GPT-5 72.6%±2.3 68.6%±4.5 80.3%±4.4 30.1%±9.0 44.2%±2.5 45.4%±1.3 73.1%±3.9 73.7%±2.7 46.6%±0.0
Every result from this document11Claude Fable 5 65.0
Printed as 65.0%Official leaderboard, measured Aug 2026MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)10Claude Fable 5Anthropic 65.0%
Every result from this document12Gemini 3.1 Pro 62.6
Printed as 62.6%Official leaderboard, measured Aug 2026MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column11Gemini 3.1 ProGoogle 62.6%
Every result from this document13Gemini 2.5 Pro 61.9
Printed as 61.9%Official leaderboard, measured Aug 2026MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)12Gemini 2.5 ProGoogle 61.9%
Every result from this document14Qwen3.5 397B A17B 61.1
Printed as 61.1%Official leaderboard, measured Aug 2026MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)13Qwen3.5 397B A17BOSSAlibaba 61.1%
Every result from this document15Kimi K2.6 59.1
Printed as 59.1%Official leaderboard, measured Aug 2026MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)14Kimi K2.6OSSMoonshot AI 59.1%
Every result from this document16GLM 5.1 First Do NOHARM v2 overall; preview; open-weight model; run date unpublished 57.9
Printed as 57.9%Official leaderboard, measured Sep 2026Configuration: First Do NOHARM v2 overall; preview; open-weight model; run date unpublishedARISE MAST technical leaderboard official leaderboard, ARISE. First Do NOHARM v2, overall leaderboard; GLM 5.1 row; displayed 2026-09-30.First Do NOHARM v2 overall metric across 19 models 15GLM 5.1OSSZ.ai 57.9%
Every result from this document17DeepSeek R1 55.8
Printed as 55.8%Official leaderboard, measured Aug 2026MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)16DeepSeek R1OSSDeepSeek 55.8%
Every result from this document
Documents
2
- ARISE MAST technical leaderboardofficial leaderboard, ARISEResults it supports
- MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results)official leaderboard, ARISE AI Research Network, 15 Aug 2026Results it supports