Clinical Benchmarks

Safety benchmarks

2 tracked · updated August 16, 2026

These measure harm avoidance rather than capability: how often recommendations carry potential for severe harm, and how models handle mental-health, self-harm, and emotional-reliance conversations. High scores here are a floor requirement for deployment, not a bragging right, and one of the two boards is a vendor's own internal evaluation, labeled as such.

First, Do NOHARM (v2)

Stanford/Harvard consortium · 1,100 consultation cases
  1. 1Anthropic logoClaude Opus 574.6%
  2. 2Moonshot AI logoKimi K374.0%
  3. 3OpenAI logoGPT-5.6 Sol70.1%
  4. 4OpenAI logoGPT-5.570.0%
  5. 5Google logoGemini 3.1 Pro62.6%
via ARISE MAST technical leaderboard · updated 2026-08full detail

OpenAI Dynamic Mental Health Evaluations

OpenAI · 3 safety metrics
  1. 1OpenAI logoGPT-5.5 Instant (June Update)0.991
  2. 2OpenAI logoGPT-5.6 Sol (August)0.981
  3. 3OpenAI logoGPT-5.6 Luna (August)0.977
via OpenAI August Updates · updated 2026-08full detail

Which safety benchmarks have current frontier-model results?

2 as of August 16, 2026: First, Do NOHARM (v2) (Claude Opus 5 leads at 74.6%); OpenAI Dynamic Mental Health Evaluations (GPT-5.5 Instant (June Update) leads at 0.991).

The other categories sit on the index: rubric-graded benchmarks, agentic and workflow benchmarks, documentation and coding benchmarks, knowledge and exam benchmarks, composite indices.