Clinical Benchmarks

First, Do NOHARM (v2)

How often, and how severely, model consultation recommendations contain potentially harmful errors.

Stanford/Harvard consortiumOfficial pagePaper
More about this boardLess

Frequency and severity of potentially harmful errors in LLM-generated medical consultation recommendations (Numerous Options Harm Assessment for Risk in Medicine); primary-care-to-specialist consults.

Paper: arXiv 2512.01241. The v1 study found potential for severe harm in up to 24.6 percent of directly applied recommendations, with errors of omission behind more than 80 percent of the severe cases.

Published by Stanford/Harvard-led consortium (50+ researchers incl. 29 board-certified physicians); hosted by ARISE, released Dec 2025. 1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options. Percentage safety score, higher better.

Headline metric
F1 (Weighted)
Board marked as preview
Yes
Board last updated
15 Aug 2026
Models in the full dataset
43
Models in the board's default view
19
Rows
17
Models
17 of 19on the board
Labs
12
Last measured
Sep 2026
Leader
86.2LiSA 2.5

Ranking

Percentage safety score, higher better

  • Anthropic
  • Google
  • OpenAI
  • Alibaba
  • Meta
  • Moonshot AI
  • Other labs
All official leaderboard
  1. 1LiSA 2.5First Do NOHARM v2 overall86.2
  2. 2Doximity Ask 6.1First Do NOHARM v2 overall84.5
  3. 3OpenEvidenceFirst Do NOHARM v2 overall80.0
  4. 4Glass 5.6 MaxFirst Do NOHARM v2 overall79.7
  5. 4Muse Spark 1.179.7
  6. 6Claude Opus 5v2 run on ARISE74.6
  7. 7Kimi K374.0
  8. 8GPT-5.6 Sol70.1
  9. 9GPT-5.570.0
  10. 10GPT-5from the Model Leaderboard SAFETY column68.6
  11. 11Claude Fable 565.0
  12. 12Gemini 3.1 Pro62.6
  13. 13Gemini 2.5 Pro61.9
  14. 14Qwen3.5 397B A17B61.1
  15. 15Kimi K2.659.1
  16. 16GLM 5.1First Do NOHARM v2 overall57.9
  17. 17DeepSeek R155.8

17 of 17 rows

Rows and sources

Open a row for the quote, the page and the document.

  1. 1LiSA 2.5 First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublished 86.2
    Printed as 86.2%Official leaderboard, measured Sep 2026Configuration: First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublished
    ARISE MAST technical leaderboard official leaderboard, ARISE. First Do NOHARM v2, overall leaderboard; LiSA 2.5 row; displayed 2026-09-30.
    First Do NOHARM v2 overall metric across 19 models 1LiSA 2.5RAGAMBOSS 86.2%
    Every result from this document
  2. 2Doximity Ask 6.1 First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublished 84.5
    Printed as 84.5%Official leaderboard, measured Sep 2026Configuration: First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublished
    ARISE MAST technical leaderboard official leaderboard, ARISE. First Do NOHARM v2, overall leaderboard; Doximity Ask 6.1 row; displayed 2026-09-30.
    First Do NOHARM v2 overall metric across 19 models 2Doximity Ask 6.1RAGDoximity 84.5%
    Every result from this document
  3. 3OpenEvidence First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublished 80.0
    Printed as 80.0%Official leaderboard, measured Sep 2026Configuration: First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublished
    ARISE MAST technical leaderboard official leaderboard, ARISE. First Do NOHARM v2, overall leaderboard; OpenEvidence row; displayed 2026-09-30.
    First Do NOHARM v2 overall metric across 19 models 3OpenEvidenceRAGOpenEvidence 80.0%
    Every result from this document
  4. 4Glass 5.6 Max First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublished 79.7
    Printed as 79.7%Official leaderboard, measured Sep 2026Configuration: First Do NOHARM v2 overall; preview; RAG clinical product; run date unpublished
    ARISE MAST technical leaderboard official leaderboard, ARISE. First Do NOHARM v2, overall leaderboard; Glass 5.6 Max row; displayed 2026-09-30.
    First Do NOHARM v2 overall metric across 19 models 4Glass 5.6 MaxRAGGlass Health 79.7%
    Every result from this document
  5. 4Muse Spark 1.1 79.7
    Printed as 79.7%Official leaderboard, measured Aug 2026
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    5Muse Spark 1.1Meta 79.7%
    Every result from this document
  6. 6Claude Opus 5 v2 run on ARISE; 19 models on the board 74.6
    Printed as 74.6%Official leaderboard, measured Aug 2026Configuration: v2 run on ARISE; 19 models on the board
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    6Claude Opus 5Anthropic 74.6%
    Every result from this document
  7. 7Kimi K3 74.0
    Printed as 74.0%Official leaderboard, measured Aug 2026
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    7Kimi K3OSSMoonshot AI 74.0%
    Every result from this document
  8. 8GPT-5.6 Sol 70.1
    Printed as 70.1%Official leaderboard, measured Aug 2026
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    8GPT-5.6 SolOpenAI 70.1%
    Every result from this document
  9. 9GPT-5.5 70.0
    Printed as 70.0%Official leaderboard, measured Aug 2026
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    9GPT-5.5OpenAI 70.0%
    Every result from this document
  10. 10GPT-5 from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI)… 68.6
    Printed as 68.6%Official leaderboard, measured Aug 2026Configuration: from the Model Leaderboard SAFETY column (NOHARM v2 F1 weighted, shown with CI); not in the Latest Flagships ranking
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, Model Leaderboard (Top 10 shown), row 10, SAFETY column
    10 GPT-5 72.6%±2.3 68.6%±4.5 80.3%±4.4 30.1%±9.0 44.2%±2.5 45.4%±1.3 73.1%±3.9 73.7%±2.7 46.6%±0.0
    Every result from this document
  11. 11Claude Fable 5 65.0
    Printed as 65.0%Official leaderboard, measured Aug 2026
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    10Claude Fable 5Anthropic 65.0%
    Every result from this document
  12. 12Gemini 3.1 Pro 62.6
    Printed as 62.6%Official leaderboard, measured Aug 2026
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view), plus Model Leaderboard SAFETY column
    11Gemini 3.1 ProGoogle 62.6%
    Every result from this document
  13. 13Gemini 2.5 Pro 61.9
    Printed as 61.9%Official leaderboard, measured Aug 2026
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    12Gemini 2.5 ProGoogle 61.9%
    Every result from this document
  14. 14Qwen3.5 397B A17B 61.1
    Printed as 61.1%Official leaderboard, measured Aug 2026
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    13Qwen3.5 397B A17BOSSAlibaba 61.1%
    Every result from this document
  15. 15Kimi K2.6 59.1
    Printed as 59.1%Official leaderboard, measured Aug 2026
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    14Kimi K2.6OSSMoonshot AI 59.1%
    Every result from this document
  16. 16GLM 5.1 First Do NOHARM v2 overall; preview; open-weight model; run date unpublished 57.9
    Printed as 57.9%Official leaderboard, measured Sep 2026Configuration: First Do NOHARM v2 overall; preview; open-weight model; run date unpublished
    ARISE MAST technical leaderboard official leaderboard, ARISE. First Do NOHARM v2, overall leaderboard; GLM 5.1 row; displayed 2026-09-30.
    First Do NOHARM v2 overall metric across 19 models 15GLM 5.1OSSZ.ai 57.9%
    Every result from this document
  17. 17DeepSeek R1 55.8
    Printed as 55.8%Official leaderboard, measured Aug 2026
    MAST technical leaderboard (First Do NOHARM v2 and per-benchmark results) official leaderboard, ARISE AI Research Network, 15 Aug 2026. arise-ai.org/mast/technical, 'First Do NOHARM v2 overall metric across 19 models' ranking (Latest Flagships view)
    16DeepSeek R1OSSDeepSeek 55.8%
    Every result from this document

Documents

2