OpenAI Dynamic Mental Health Evaluations: current results
OpenAI · dynamic simulated conversations (counts not disclosed) · index updated August 16, 2026
GPT-5.5 Instant (June Update) holds the top current result on OpenAI Dynamic Mental Health Evaluations, 0.991 as of 2026-08, per OpenAI GPT-5.6 August Updates (PDF). Multi-turn adversarial user simulations for mental health, emotional reliance, and self-harm response quality, where conversations evolve in response to model outputs rather than following fixed scripts.
Current results
- 1
GPT-5.5 Instant (June Update)0.991
- 2
GPT-5.6 Sol (August)0.981
- 3
GPT-5.6 Luna (August)0.977
Result detail
| # | model | score | as of | |
|---|---|---|---|---|
| 1 | GPT-5.5 Instant (June Update) OpenAI mental health 0.991, emotional reliance 0.989, self-harm 0.967; measured at lowest reasoning deployment settings | 0.991 | 2026-08 | |
| 2 | GPT-5.6 Sol (August) OpenAI mental health 0.981, emotional reliance 0.961, self-harm 0.901; OpenAI flags statistically significant offline self-harm regression vs GPT-5.5 June, not reproduced online | 0.981 | 2026-08 | |
| 3 | GPT-5.6 Luna (August) OpenAI mental health 0.977, emotional reliance 0.965, self-harm 0.911 | 0.977 | 2026-08 | |
Scores appear exactly as OpenAI GPT-5.6 August Updates (PDF) publishes them (vendor-reported scores). An internal OpenAI safety evaluation covering OpenAI models only; it is not independently runnable, and OpenAI notes the error rates are not representative of average production traffic. Listed as a vendor safety eval, not a cross-vendor benchmark.
About the benchmark
| publisher | OpenAI |
|---|---|
| category | safety benchmarks |
| released | 2026-08 |
| size | dynamic simulated conversations (counts not disclosed) |
| scale | compliance rate per metric, 0 to 1, higher better; headline number is the mental-health metric |
| result basis | vendor-reported scores |
| source | OpenAI GPT-5.6 August Updates (PDF) |
| last frontier result | 2026-08 |
What is OpenAI Dynamic Mental Health Evaluations?
OpenAI Dynamic Mental Health Evaluations is a safety benchmark from OpenAI, released 2026-08: dynamic simulated conversations (counts not disclosed), scored on a compliance rate per metric scale. Multi-turn adversarial user simulations for mental health, emotional reliance, and self-harm response quality, where conversations evolve in response to model outputs rather than following fixed scripts.
Which model leads OpenAI Dynamic Mental Health Evaluations?
GPT-5.5 Instant (June Update) (OpenAI) holds the top current result on OpenAI Dynamic Mental Health Evaluations at 0.991, per OpenAI GPT-5.6 August Updates (PDF), as of 2026-08.
Where do the OpenAI Dynamic Mental Health Evaluations numbers come from?
From OpenAI GPT-5.6 August Updates (PDF) (vendor-reported scores). An internal OpenAI safety evaluation covering OpenAI models only; it is not independently runnable, and OpenAI notes the error rates are not representative of average production traffic. Listed as a vendor safety eval, not a cross-vendor benchmark.
The rest of the field is on the index, and how sources qualify is on the methodology page. Model names in the table link to cross-benchmark pages.