Clinical Benchmarks

First, Do NOHARM (v2): current results

Stanford/Harvard-led consortium (50+ researchers incl. 29 board-certified physicians); hosted by ARISE · 1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options · index updated August 16, 2026

Claude Opus 5 holds the top current result on First, Do NOHARM (v2), 74.6% as of 2026-08, per ARISE MAST technical leaderboard. Frequency and severity of potentially harmful errors in LLM-generated medical consultation recommendations (Numerous Options Harm Assessment for Risk in Medicine); primary-care-to-specialist consults.

Current results

Result detail

#modelscoreas of
1Anthropic logoClaude Opus 5 Anthropic
v2 run on ARISE; 19 models on the board
74.6%2026-08
2Moonshot AI logoKimi K3 Moonshot AI74.0%2026-08
3OpenAI logoGPT-5.6 Sol OpenAI70.1%2026-08
4OpenAI logoGPT-5.5 OpenAI70.0%2026-08
5Google logoGemini 3.1 Pro Google62.6%2026-08

Scores appear exactly as ARISE MAST technical leaderboard publishes them (official leaderboard). Paper: arXiv 2512.01241. The v1 study found potential for severe harm in up to 24.6 percent of directly applied recommendations, with errors of omission behind more than 80 percent of the severe cases.

About the benchmark

publisherStanford/Harvard-led consortium (50+ researchers incl. 29 board-certified physicians); hosted by ARISE
categorysafety benchmarks
released2025-12
size1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options
scalepercentage safety score, higher better
result basisofficial leaderboard
sourceARISE MAST technical leaderboard
last frontier result2026-08

What is First, Do NOHARM (v2)?

First, Do NOHARM (v2) is a safety benchmark from Stanford/Harvard consortium, released 2025-12: 1,100 consultation cases, 10 specialties, 12,747 expert annotations on 4,249 management options, scored on a percentage safety score scale. Frequency and severity of potentially harmful errors in LLM-generated medical consultation recommendations (Numerous Options Harm Assessment for Risk in Medicine); primary-care-to-specialist consults.

Which model leads First, Do NOHARM (v2)?

Claude Opus 5 (Anthropic) holds the top current result on First, Do NOHARM (v2) at 74.6%, per ARISE MAST technical leaderboard, as of 2026-08.

Where do the First, Do NOHARM (v2) numbers come from?

From ARISE MAST technical leaderboard (official leaderboard). Paper: arXiv 2512.01241. The v1 study found potential for severe harm in up to 24.6 percent of directly applied recommendations, with errors of omission behind more than 80 percent of the severe cases.

The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.